WTKRESEARCH + ENGINEERING
← All engineering journal records

Research Notes

A Judge Can Only Ground What the Evidence Makes Answerable

Better judges, better holdouts, and better repairs all depend on knowing exactly what evidence is available.

A new personalization study challenges a convenient version of one of our own assumptions: that fabrication cannot be judged reliably from answer text. The more useful lesson is narrower. A judge can evaluate a claim when the evidence it is allowed to use is complete enough to make the question answerable.

The paper, The Personalization Mirage, studied models that invent user attributes beyond the available evidence. Its independent judge was compared with a blind human annotator on 400 claims and reached Cohen's kappa of 0.863 across four classes and 0.900 for the binary result. That is meaningful agreement, not a casual model preference.

There is an important boundary. The benchmark uses 150 constructed personas. The admissible evidence is closed and enumerable, so a judge can compare each claim with a known source of truth. That is not the same as deciding whether an open-world answer has invented something somewhere in a large, incomplete body of evidence.

WTK research should state the distinction more precisely. The question is not whether a claim appears in answer text. The question is whether the evidence needed to assess that claim is available, attributable, and bounded. When it is not, the honest result is not assessable, not pass.

The paper also reports an exploratory warning about self-monitoring. Across 12 models, self-assessed over-inference was negatively rank-correlated with judged over-inference. The bootstrap confidence interval runs from -0.90 to +0.06, so we should not present the inversion as settled. We should still avoid using a model's confidence or self-critique as an honesty verdict across models. The external evidence path matters more than how confident the model sounds.

Also on the radar

Held-out should mean attributable. GDPevo decomposes business workflows into atomic rules, distributes subsets across training tasks, then recombines them in held-out tests. That construction helps distinguish a reusable improvement from memorizing a fixture. The paper reports held-out accuracy gains of up to 16.44 percentage points, while the best evolved agents remain below a 91.6 percent oracle ceiling. Those are value results, not honesty or safety evidence. For WTK, the useful idea is the experiment design, not another benchmark to adopt.

A repair needs an operator and a retest. SKILL-KD compares a failed student trajectory with a teacher trajectory, derives a textual skill patch, and reruns the frozen student. Its repair vocabulary is concrete: add, delete, modify, or skip. That is useful for bounded improvement, but delete and modify are also the dangerous operators. They must never weaken a must-pass validity rule, and a candidate repair must qualify again while the incumbent remains available.

What WTK should test

These papers point to one shared experiment discipline. First declare what evidence makes each obligation answerable. Then construct held-out cases that separate reusable behavior from memorization. Finally, test a bounded repair without allowing it to change the definition of success.

Still unknown is where the boundary between assessable and open-world claims should sit for each goal class. That is the next useful test. Give the same judge closed evidence, incomplete evidence, and evidence that contains a known contradiction. Measure correct passes, correct failures, and abstentions separately. A judge that knows when it cannot answer may be more valuable than a larger panel that always does.

RECORD DETAILSReference RN-008
Artifact
Research Notes
Status
Published
Published
August 6, 2026
Linked sources
3