An Auditor Remembers What You Fixed
A verifier's repair history can change its threshold before the next task even begins.
If a verifier sees the same task twice, we want the same evidence to mean the same thing. A new preprint shows why that assumption needs a test: putting an earlier audit-and-repair episode in an LLM verifier's context shifted its decision threshold even when the current task was byte-identical.
The signal
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency evaluates LLM verifiers on human-verified-correct ProcessBench traces. The authors compared a prior audit-and-repair episode with a length-matched non-audit control while holding the present task fixed. In all 15 model and wording combinations, the audit-and-repair context lowered false alarms by 2.8 to 11.5 percentage points, or a 9% to 25% relative reduction.
The result is not simply a story about a model becoming less accurate. Signal-detection analysis found the reported change in decision criterion: the criterion moved in all 15 combinations and survived correction in 13, while the reported sensitivity measure, d-prime, survived in none. A hand audit found 41 of 50 false alarms, or 82%, were simply wrong, so fewer alarms may be helpful at this operating point.
The evidence
The useful lesson is about evaluator inputs. A completed repair episode can influence how a verifier interprets later evidence, even when the current task has not changed. That means a post-repair verdict can reflect both the current artifact and the verifier-visible history that preceded it.
For WTK, we should treat repair context as a declared evaluator condition, not harmless background. We already separate evaluator authority from the package under review and bind evidence to an exact artifact. Our take is that the verifier's prompt, model binding, prior context, and evaluation budget also need enough identity to tell whether a changed verdict belongs to a changed package or a changed decision rule.
The boundary
This is a preprint on selected models, author-defined prompts, and human-verified-correct ProcessBench traces. It does not test WTK packages, target projections, qualification records, or semantic evaluators. Crucially, the abstract does not report an equivalent false-negative measurement on incorrect traces, so it cannot tell us whether the apparent leniency would be safe or unsafe for a WTK gate.
It is not a WTK Finding and does not promote the maturity of any WTK mechanism. A context-controlled experiment is still needed before we make a claim about our own evaluators.
The builder impact
Builders should record what a verifier can see before it issues an acceptance-critical verdict. That includes prior audit text, repair proposals, prior verdicts, prompt wording, model settings, tools, and token limits. If repair history is allowed in, make it an intentional arm of the evaluation, not a residue from the previous run.
The WTK test
WTK could run a small held-out evaluator-context comparison. Freeze the package, positive and negative observations, evaluator authority, prompt, model binding, sampling, rubric, and budget. Present each observation under a clean context and under a sealed prior audit-and-repair context, retain every verdict, and measure both false-positive and false-negative changes.
That would give us a bounded answer about one evaluator and execution form. Until then, this paper is a reason to control verifier history, not evidence that WTK's gates have drifted.
Have an approach, result, or counterexample?
You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.
Contribute to this research question →Working with an AI assistant?
Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.
Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.