WTKRESEARCH + ENGINEERING
← All engineering journal records

Research Notes

"Held-Out" Doesn't Mean Clean

In one multimodal fact-checking study, up to 29% of post-cutoff claims remained potentially contaminated, and the effect changed system rankings. That challenges how comparative value is measured, including in WTK.

Wednesday's lead catch goes after an assumption most evaluation work quietly stands on, and WTK stands on it too, so it gets the deep dive.

The catch: your fresh test set isn't as fresh as you think

The standard defense against a model having memorized your benchmark is to build the test set from material published after the model's knowledge cutoff. A new study (Novel Claim or Deja Vu?) measures how well that actually works in fact-checking evals. The findings: dynamic evaluation reduces contamination but doesn't eliminate it, with 17.09% to 29.30% of post-cutoff claims remaining potentially contaminated. Many "new" claims are resolvable by stitching together pre-cutoff knowledge. And the punchline: contamination inflated scores by up to 11.34 points of Macro-F1 and distorted system rankings.

That last one is the bite. WTK's quality bar is defined as measured held-out, and our core value claim is comparative: a WTK-composed agent has to beat the incumbent it would replace, which is a ranking. If contamination can reorder rankings by that much, then "held-out by publication date" alone isn't enough to make a structure-adds-value claim safe. Part of a measured lift could be differential contamination rather than composition doing real work.

How WTK research proposes to handle it

Honest scope first: the paper measured multimodal fact-checking, not goal-accomplishment tasks. It challenges the measurement assumption; it is not evidence that any specific WTK number is wrong.

Our response is declarative, not mechanical. We're not building a contamination detector. Instead, the definition of "held-out" that our quality numbers rely on gets a stated contamination-control clause, so any published number carries a disclosure of how its test set was constructed and controlled. Cheap to state now, embarrassing to retrofit after someone asks. This fits a theme that kept building all week: a quality number should ship wearing its full scope statement.

Also on the radar

  • The memory that doesn't look like the question. Keep It InMind names an assumption "so natural it is rarely stated": that a memory you need will resemble the query that needs it. Their counterexample is a tree-nut allergy that should change the answer to a macaron request via almond flour, sharing zero surface cues a retriever can see. With the decisive fact placed in context, the model answers 84% of indirect queries; when the same fact must be retrieved, six different memory systems top out at 14.4%. This is evidence of a query-conditioned retrieval failure mode relevant to WTK; it does not validate goal-derived provisioning as the solution. It motivates a WTK test: compare declaration-based provisioning, similarity retrieval, and hybrid routing on held-out goal tasks. It also handed us a homework item: confirm whether any similarity-shaped selection in WTK could inherit this blind spot. We also pocketed their clean three-way fault attribution: was the fact never stored, did the model lack the bridging knowledge, or was it stored and never surfaced?
  • One benchmark result for text-level fabrication detection. A span-level hallucination detection paper (Beyond Document Grounding) built a detector over code, tool output, and documents: its fine-tuned model reaches 0.689 span-F1, and 0.60 on code-agent output, while the reported zero-shot LLM judges top out at 0.22. These results do not establish a theoretical ceiling and do not justify calling any finite test result "honesty." They support keeping deterministic claim checks separate from probabilistic detectors and semantic judges. A blocking check may require zero failures on its declared test set, but that is not a claim of complete safety or honesty.

The takeaway

Evaluation is where grounding and claim-support assertions get tested, and this week's lesson is that the test itself needs receipts. Say how your held-out set was built. Say what your detector can't catch. WTK's whole posture is that a number without its scope statement is a marketing number, and now we have two more citations for why.

See you at the next catch.

RECORD DETAILSReference RN-003
Artifact
Research Notes
Status
Published
Published
July 29, 2026
Linked sources
3