A Judge Cannot Promote Its Own Homework
A label made by an improvement loop is still a hypothesis until independent evidence checks it.
An improvement loop can produce a tidy scorecard long before it has earned the right to trust its labels. For an AI agent that proposes, judges, and promotes its own changes, that gap is not cosmetic. It determines whether a better internal score is evidence of progress or merely a story the loop told itself.
For example, a coding agent might prefer a patch because it was assembled by specialists rather than written in one pass. That is a useful candidate to test, not permission to deploy it without independent checks.
The signal
J-Zero trains a Challenger, Solver, and Judge together using preference pairs whose order is assumed from the way the responses were made. One pair type treats a Solver answer as better than a Challenger answer. The other treats a response assembled from subtask answers as better than the Solver's one-shot response.
That second assumption failed its own early audit. The authors had an external LLM judge compare both response orders, excluding ties. The supposed subtask-amplification winner was chosen only 21.1% of the time in iteration one and stayed below 50% for the first three iterations. It crossed 50% only from iteration four and later reached roughly 70 to 80%. The early mismatch matters, but so does that recovery: the audit does not show that decomposition is uniformly worse.
The evidence
The other label class held up better in the same check: the Solver response won more than 60% of comparisons at every iteration, though its rate declined from 87.9% to about 66%. Under the paper's selected benchmarks and checkpoint rule, J-Zero raised the Qwen3-4B base model's reported overall average from 44.91 to 54.38 on verifiable tasks, and from 9.58 to 20.81 on unverifiable tasks.
Those results make the useful contradiction sharper, not weaker. The system can show better benchmark scores while one source of its evaluator's training labels begins as a bad ordering. A rising internal score does not settle whether each kind of training signal was sound enough to influence the candidate.
The boundary
This is a preprint about the authors' Qwen base models, Skywork reward model, benchmarks, training procedure, and best-checkpoint rule. Larger and post-trained reasoning models were not tested as the adapting Solver. Its label audit uses an external model judge, not human ground truth, and drops ties. It does not evaluate WTK packages, compiled targets, independent operator authority, or real deployed goals.
We should not conclude that evaluator adaptation is always unsafe, that the reported gains are false, or that decomposition is useless. The paper shows a specific failure boundary: production history alone does not make a preference ordering trustworthy.
The builder impact
WTK research views this as a reason to separate candidate generation from promotion evidence. A role name, a task decomposition, a score increase, or a same-loop judge may suggest a hypothesis. None should become the acceptance authority for a changed agent or harness.
WTK already has the right direction in its claim-to-evidence lifecycle and research and falsification method: retain the evidence, keep a held-out check separate, and make promotion an independently governed decision. The paper gives that boundary a useful failure-shaped test.
Our earlier note on independent reviewers and their evidence focused on who can review and what they can inspect. Here the narrower question is whether each label class deserves to train the evaluator at all.
The WTK test
Our proposed test freezes a baseline package, target, evaluator, and a one-use held-out goal partition before any evaluator adaptation begins. That means reserving tasks the adapting system cannot see and using them once for the final comparison against the unchanged baseline. For every self-generated preference class, compare its ordering with labels from an independently authorized review path, including ties and abstentions. A class cannot influence adaptation until it clears a predeclared agreement threshold.
Retain every candidate, label source, evaluator version, checkpoint decision, held-out receipt, rejection, and final target outcome. The test fails if self-generated labels do not meet the threshold, if an adaptive component can see held-out goals, or if a candidate's target-bound result does not improve without weakening a control. WTK has not run this test. External research has given us a hypothesis, not a maturity promotion. Any change to the model, evaluator, target harness, or label policy would require a fresh comparison.
Source
Primary research: J-Zero: Unified Challenger-Solver-Judge Co-Evolution from Zero Data, arXiv:2608.26582. WTK reviewed the canonical 19-page paper and treats it as external evidence for a proposed test, not evidence that WTK already passes that test.
Have an approach, result, or counterexample?
You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.
Contribute to this research question →Working with an AI assistant?
Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.
Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.