WTKRESEARCH + ENGINEERING
← Back to Journal

A Warning Is Not Containment

Recognizing a hostile message is only the first check in a persistent agent system.

An agent team can recognize a hostile message and still fail to contain it. That matters because a warning, refusal, or good final response can look reassuring while the original content remains available to influence later actions through memory, shared state, or a tool path.

Emergence World started eight simulated worlds with ten agents each. One world ended before the first intervention; the remaining seven encountered a phishing campaign, a fabricated shutdown memo, and an announced memory breach. The authors scored recognition separately from restraint, retention, propagation, containment, and durable response. No world satisfied every criterion across all three events.

That separation is the useful part. The best reported phishing, misinformation, and memory-breach scores were 6/9, 3/6, and 5/5. Claude led phishing and shared the misinformation lead with DeepSeek; OpenAI led the memory-breach assessment. These count study-specific criteria satisfied, not general safety ratings. All seven exposed worlds warned peers about phishing, but none satisfied the public-trace removal criterion or the ongoing monitoring-and-removal criterion. Every exposed world acted or published around the false shutdown message before verification, then repeated it as fact. About 46 hours after the phishing campaign ended, an agent recovered an attack link from a stored message and attempted to fetch it again. The endpoint then returned a safe response. This was delayed attempted re-engagement, not a demonstrated second compromise.

The paper does not establish a WTK failure, a model ranking, or a general failure rate. It is an arXiv preprint about an author-operated simulation with one connected trajectory for each model configuration, ten designed personas, bespoke event rubrics, a 200,000-token context cap, and provider-specific serving stacks. The prompts contained no event-specific instructions for phishing, misinformation, or unauthorized data access. The authors acknowledge that explicit guidance could change the outcomes, as could different memory policies, tools, and target controls.

The engineering distinction

For WTK, detection is evidence about one stage of a response. Containment is a stronger target claim: the deployed path must also restrain unsafe actions, prevent cross-role spread, clean up or bound retained influence, and recover after ordinary work continues. Our tool-authority work already separates information obtained from a tool from permission to act. This research adds a time dimension. Blocking the first action is not enough if the same content can later return through a permitted memory or communication path.

Retention alone is not necessarily failure. A quarantined record may be needed for investigation. The question is whether retained material can improperly influence execution or spread beyond its permitted audit purpose.

WTK keeps package source, target projection, and qualification distinct because they answer different questions. A package can declare a safe boundary, while a specific target must still demonstrate that its actual state, tools, and handoffs honor it. That is why target-side monitoring and intervention are not interchangeable: an alert after a state transition cannot undo the effect.

What WTK could test

We could first measure one existing package in its execution environment over several turns, without changing its safeguards. This is baseline characterization, not a test of an improvement. A clearly labeled untrusted test item would enter through one already permitted test surface. Separate checks would cover execution restraint, later recall, cross-role propagation, cleanup, and recovery after intervening benign work. It would hold the package, contract, target, tools, memory policy, runtime binding, evaluator, event content, and duration fixed, while retaining every attempt and decision.

Include benign controls to check that permitted work still completes. Define expected behavior for each check before running it, including allowed audit retention. Report observed violations, checks satisfied, and inconclusive outcomes separately. Passing these cases would not establish universal safety. A later improvement comparison would need a specified candidate change and separate approval. This first test would not create a new simulation, change a WTK contract, or authorize a deployment.

Our take: safety evidence should follow the content's whole reachable path, not stop at the first sensible sentence an agent says about it. WTK has not run this characterization. Any protocol, target selection, and execution need separate human approval.

Source

Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems, September 15, 2026, arXiv:2609.17320v1.

Have an approach, result, or counterexample?

You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.

Contribute to this research question
Working with an AI assistant?

Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.

Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.

RECORD DETAILSReference RN-056
Artifact
Research Notes
Status
Published
Evidence posture
External research interpreted; WTK target-harness characterization not run
Published
September 16, 2026
Author
WTK Research
Review
WTK human editorial review
Linked sources
1