WTKRESEARCH + ENGINEERING
← Back to Journal

A Retrieval Refresh Can Change the Answer

A green utility score can still hide a changed answer regime.

A retrieval refresh can leave the model name, prompt, and aggregate score nearly unchanged while changing which answer an agent gives. That is the useful contradiction in a new study of retrieval-augmented question answering: a dashboard can look calm while the answers underneath it have moved.

The signal

Same Agent, Different Answers compares a fixed single-turn retriever-generator at two versions of its accessible corpus. The authors held the requested model identifier, prompt, retriever, top-eight evidence depth, rendering, and exposed generation controls fixed. They then took two independent answers per question at each corpus snapshot, so ordinary repeat variation did not masquerade as an update effect.

On a preregistered 400-question Natural Questions study, growing one nested FineWeb path from one shard to seven produced 6.438 percentage points of normalized-exact excess churn and 10.250 points of blind-semantic excess churn. Exact-match accuracy moved only minus 1.50 points. In other words, gains and losses can largely cancel in an aggregate score while individual answers no longer behave as they did before.

The evidence

The paper calls its method a Snapshot Compatibility Audit. It subtracts disagreement between repeated calls to the same snapshot from disagreement across the two snapshots. That distinction matters. A one-shot before-and-after comparison would count normal generation variability as corpus-induced change.

The main Natural Questions result cleared the study's preregistered gates, with one-sided 95 percent lower bounds of 4.563 points for exact churn and 7.688 points for semantic churn. A separate 200-question TriviaQA study was supportive, not pooled, and reported 3.000 and 2.125 points while exact match increased 1.25 points. An outcome-blind 100-question rerun with a second DeepSeek generator and serving configuration also found positive churn while exact match rose.

The boundary

This is a 2026 preprint about a single-turn QA system, one nested FineWeb path, one search service, top-eight retrieval, two English benchmarks, and two DeepSeek generator and serving configurations. It does not test a WTK package, a planning agent, multi-turn tool use, or a memory-writing system. It also does not show that every changed answer is harmful, identify the document responsible for a change, or set a universal release threshold.

Two repeats provide only a coarse view of each answer distribution. The semantic-label audit used a second judge family, not human validation. The reported intervals describe the frozen question cohorts, not alternative corpus partitions, prompts, models, or deployment dates. This is a mechanism signal, not a current WTK performance claim.

The builder impact

For WTK, package identity and target qualification remain necessary, but they cannot make a retrieval or evidence-snapshot update behaviorally neutral by themselves. If an agent's grounded answer depends on an external evidence surface, that surface needs an explicit version and a compatibility question beside the usual utility question.

Our take: treat a material retrieval refresh as a possible behavior change even when the package source, model binding, API, and average evaluation all look unchanged. The grounding floor is about whether evidence is attributable and usable. Qualification should also be able to ask whether a controlled evidence change moved a sensitive answer beyond ordinary repeat noise.

The WTK test

WTK could freeze a package digest, model binding, target harness, held-out goals, retrieval settings, and old and new evidence snapshots. For every goal and snapshot, collect at least two independent target outputs, retain every attempt and evidence receipt, and report same-snapshot agreement, cross-snapshot agreement, excess churn, and the existing utility measure. Blind review of high-impact flips should distinguish a helpful correction from an unsafe or unexplained change.

Predeclare the decision threshold for this target and corpus transition. The test fails if it hides changed outputs behind an average score, drops failed attempts, or cannot bind each result to both the package and evidence snapshot. Passing would not make retrieval updates universally safe. It would show that one declared update survived a compatibility check designed to reveal a quiet answer shift.

Have an approach, result, or counterexample?

You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.

Contribute to this research question
Working with an AI assistant?

Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.

Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.

RECORD DETAILSReference RN-034
Artifact
Research Notes
Status
Published
Evidence posture
Published with the evidence boundary stated in this record
Published
August 25, 2026
Author
WTK Research
Review
WTK human editorial review
Linked sources
1