A Precise Score Can Still Measure the Wrong Thing
A score only speaks for the behavior its cases can actually distinguish.
An evaluation can be perfectly scored and still fail to measure the behavior named in its claim. That matters for builders because a clean, repeatable score is not enough if materially different policies would look identical under the cases we chose to observe.
The signal
Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation asks a sharper question than whether an evaluator is implemented correctly: can its observation support distinguish the policies that differ on the property it claims to measure?
In the authors' finite, solver-grounded diagnostic, base-only observation collapsed seven frozen deterministic policies into one equivalence class. Full observation separated all seven. Every leave-one-out version retained a constructive collision witness, meaning a missing observation left at least two different policies looking the same on the claimed behavior.
The evidence
The paper also compares two constrained-generation variants. Both achieved pair-validity of 1.0, yet base accuracy and selective-response fidelity diverged: 0.620 versus 0.324 across six balanced oracle-transition directions, with the reported gap recurring on a second deterministic source.
That distinction is useful because it separates a familiar question, "Did the answers match here?", from a harder one, "Could these cases reveal the behavioral difference we say we care about?" The authors synthesize a two-cell identifying support for their frozen policy class from a 36-cell tensor without making more model calls.
The boundary
This is a controlled diagnostic over a finite policy class. It does not show that WTK's current qualification fixtures are non-identifying, that every agent behavior has a small identifying test set, or that symbolic analysis can replace held-out evaluation and target evidence.
It also does not establish that a deterministic check is weak. A deterministic check can be valuable while still being unable to support a broader claim than its observations can distinguish.
The builder impact
Our take: every qualification claim needs a validity question before it needs another score. Name the behavior alternatives that matter, then ask whether the declared cases and oracle would separate them. If two materially different outcomes collapse to the same passing record, the score is evidence of a local result, not evidence for the wider behavior.
WTK can keep this narrow. Package, target, and evaluation identities still matter. Retaining every attempt still matters. The extra discipline is to state what a fixture is supposed to distinguish, and to preserve a collision witness when it cannot do so.
The WTK test
For a bounded WTK qualification claim, define the behavior alternatives and observable consequences before inference. Exercise the declared cases against those alternatives, retain any collision witnesses, and label the resulting score with the property the protocol can actually identify.
That would test the adequacy of one declared evaluation protocol. It would not prove general reliability, replace independent review, or promote a package because one fixture passed.
Still unknown
We do not yet know which WTK qualification claims have identifying support across different goals, targets, and harnesses. The next useful test is a predeclared fixture audit that compares behaviorally different cases under one stated claim, retains every result, and reports whether the support separates the difference it is meant to measure.
Have an approach, result, or counterexample?
You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.
Contribute to this research question →Working with an AI assistant?
Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.
Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.