WTKRESEARCH + ENGINEERING
← Back to Journal

Relevant Memory Can Still Be the Wrong Context

Correct history can still steer an agent away from the goal it has now.

A memory can be factually right and still nudge an agent toward the wrong answer. That is the useful warning in a new benchmark: relevance, provenance, and past success are not enough to show that a remembered context helps the goal in front of the agent.

The signal

MemTrapBench studies cases where historical context leads a model to carry an earlier strategy, boundary, or belief into a task whose conditions have changed. The authors constructed 1,050 examples across cognitive bias, task boundary, trauma, and safety scenarios, then compared two model families with five memory strategies against a no-memory baseline.

The contradiction is worth holding onto. A memory can be semantically related and valid in its original setting, yet still make the current answer worse because it narrows the strategy the model considers or extends an old exception past its proper scope.

The evidence

In the paper's aggregate results, the no-memory baselines scored 85.16 percent for Gemini-3-Flash-Preview and 81.83 percent for Qwen3-30B-A3B-Instruct-2507. Every evaluated memory strategy scored lower. The strongest Gemini result was 71.17 percent for EverMemOS, and the strongest Qwen result was 70.13 percent for LightMem.

The more revealing control holds the Task Boundary query family constant. No memory scored 92.29 percent. Relevant history without the designed trap scored 94.39 percent. Trap-inducing history scored 31.05 percent. That comparison is a better builder question than "does memory help?" It asks whether this admitted history helps this task, under these conditions, compared with the alternatives.

The boundary

This is a preprint benchmark with deliberately constructed traps, two named models, and five named memory frameworks. It does not measure WTK packages, target projections, tool receipts, or live user outcomes. It also does not estimate how often harmful retrieval occurs in open work, prove that every retained context is unsafe, or validate the paper's prompt-based mitigation as a WTK control.

The study reports a useful failure mode, not a universal diagnosis. A correct source and a high retrieval score can still be valuable signals. They are simply not proof that the memory should shape a current decision.

The builder impact

Builders should separate two questions that are easy to blur: is this history trustworthy, and does it improve the current goal? The first is about source, authority, and retention. The second is counterfactual. It needs a comparison against no admitted history and against an equally trustworthy history that carries the wrong task boundary.

For WTK, that reinforces the grounding floor. Evidence linked to a prior task should not silently become authority for a new one just because it is easy to retrieve. Our take is that a memory admission path needs to preserve the old context and make its claimed relevance testable against the current obligation.

The WTK test

WTK could run a small held-out three-arm evaluation: no admitted history, provenance-correct history that helps the stated goal, and provenance-correct history that conflicts with the goal boundary. Hold the package, target harness, tools, model binding, goal, and evaluator fixed. Predeclare receipt use plus any required refusal or escalation behavior, retain every attempt, and bind the outcome to the exact package and runtime identities.

That would let us measure whether an admitted memory improves this execution form without turning retrieval into a proxy for success. Until then, this paper is a reason to test memory admission, not evidence that WTK has a memory failure.

Have an approach, result, or counterexample?

You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.

Contribute to this research question
Working with an AI assistant?

Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.

Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.

RECORD DETAILSReference RN-028
Artifact
Research Notes
Status
Published
Evidence posture
Published with the evidence boundary stated in this record
Published
August 21, 2026
Author
WTK Research
Review
WTK human editorial review
Linked sources
1