A Clean Receipt Does Not Make a Stable Judge
A hash can prove the same request was sent, not that a shared evaluator will make the same decision.
An evaluator can produce impeccable receipts and still be too unstable to carry a qualification decision. That matters when an agent's promotion, release, or tool boundary depends on an evaluator verdict rather than a deterministic check.
Imagine a reviewer that ranks four candidate fixes. The request hash, model label, and event log can all match on Tuesday and Wednesday. If the reviewer changes the ranking anyway, the record proves delivery, not measurement reliability.
The signal
Clean Engineering, Unstable Measurement treats a black-box LLM judge as an instrument whose stability must be measured. In its preregistered next-day replay, all 100 request bodies were byte identical, their hashes and visible metadata were constant, and all requests were delivered. Exact full-ranking agreement was still 78 out of 100 against a gate frozen at 0.99.
The sharp point is not that logging failed. The paper's audit trail did the useful thing: it preserved the 22 disagreements instead of changing the metric or denominator after seeing them. The failed gate describes the instrument's behavior under that setup, not a defect in the request record.
The evidence
In a separate 1,000-call battery, the study reported median exact-ranking agreement of 0.805 within a day and 0.800 across days. Its four separately reported provider arms ranged from 0.740 to 0.879 for the same structured ranking battery. Those are stability observations, not accuracy rankings or a provider leaderboard.
WTK research reads this as a boundary on evaluator identity. A receipt, a constant model name, and a valid execution audit can establish provenance. They do not establish that an evaluator is stable at the decision granularity a gate consumes. That is why the evidence lifecycle needs to bind the instrument as carefully as the artifact it assesses.
The boundary
This preprint studies one observer family on a shared endpoint, synthetic exact-rational tasks, one prompt template, and a structured-ranking readout. It does not evaluate WTK, estimate production disagreement rates, or show that any evaluator should be rejected. The four-provider comparison uses different capability tiers and does not compare judging accuracy.
The paper also does not make a local evaluator deterministic. Its useful mechanism signal is narrower: a gate written before execution is meaningful only when the measurement behavior it relies on has also been measured. WTK has not replicated this result, and a changed evaluator, target, decision set, or serving context would need a fresh characterization.
The builder impact
For WTK, model and provider settings are runtime bindings, not a complete evaluator identity. When an evaluator verdict is load-bearing, we should retain the evaluator configuration, repeated receipts, invalid-output handling, and the agreement distribution alongside the qualification packet. Deterministic conformance checks remain useful, but an LLM verdict needs its own evidence boundary.
This extends the point that a portable package does not qualify its runtime. Package and contract identity can travel cleanly while an evaluator's behavior remains specific to the target context where it is asked to judge.
The WTK test
We propose a bounded characterization study for one existing qualification profile where an evaluator verdict matters. Hold the package and contract digests, target, evaluator configuration, fixtures, permissions, model binding, and budget fixed. Compare the receipt-only profile with the same profile plus a separate evaluator-characterization receipt before its verdict can carry the gate.
Run an identical decision set in one window and a second window. Retain every request and response receipt, response-side identity signal, invalid output, ranking or verdict, and gate outcome. Measure exact agreement, disagreement shape, false blocks and false approvals against a deterministic oracle or human-reviewed fixture where available, then measure qualification outcome, cost, and latency.
The candidate fails if it cannot meet a predeclared stability rule, if it hides disagreements, if the proposed receipt changes no meaningful decision, or if it weakens deterministic policy enforcement. This is a proposed WTK experiment, not evidence that WTK currently passes the test or authorization to run it.
Source
Primary research: Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints, arXiv:2609.04198, September 3, 2026. WTK reviewed the canonical 35-page v1 paper and treats it as external evidence for a proposed evaluator-characterization test.
Have an approach, result, or counterexample?
You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.
Contribute to this research question →Working with an AI assistant?
Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.
Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.