WTKRESEARCH + ENGINEERING
← Back to Journal

A Score Gain Needs a Frozen Twin

Repeated tests of an unchanged system help distinguish measured improvement from score variation.

Editorial clarification: September 8, 2026

The summary now explains the unchanged control directly. Reported measurements, statistical limitations, and the proposed-only WTK test remain unchanged in scope.

An agent can look better after a change even when the change did nothing. The self-improvement audit discussed here shows that some measurements report gains, losses, or new capability when an unchanged model is evaluated twice under the same procedure.

The signal

Phantom Gains: Auditing Self-Improvement Against a Measured Null evaluates three rounds of rank-32 LoRA self-training on Qwen3-8B alongside a frozen, untrained control. The control goes through the same evaluation schedule as the trained arms. That turns a familiar baseline into a measured null: what does this metric say when no learning occurred?

One single greedy decode on MATH-500 made the frozen model appear to gain six problems and corrupt nine. On AIME, a one-success expansion rule assigned that same frozen model an expansion rate of 0.280. Raising the threshold to two successes looked like a repair in one comparison, but across 110 frozen comparison pairs its null was still 0.058, with a 95 percent interval of [0.038, 0.078].

The useful comparison is with an unchanged system. Treat a claimed transition as evidence only after comparing it with the transitions an unchanged system produces under the same harness, sampling, scoring, and schedule. The paper also uses a per-problem test against pooled base evaluations, which returned zero detections on its held-out frozen replicates. That is a result about this study's setup, not a general recipe guaranteed to fit every evaluation.

The evidence

The authors audited a Qwen3-8B backbone across public math benchmarks, three self-training forms, and an external-distillation control. On 22 AIME problems the base model reached at most five times in 1,408 pooled draws, distillation improved 8 to 11 per seed while the self-training arms improved 0 to 2. On the ten problems the base model never reached, the 5-versus-2 comparison was not significant. The study correctly keeps those claims separate.

The distinction matters when interpreting an improvement score. A positive transition count can be real and still be too close to the process's own measurement noise to support an improvement claim. A final score is not automatically safer: the paper shows that a fixed token cap, seed variation, and underpowered probes can each change the story told by a score.

The boundary

This is a preprint under review, not a WTK experiment. Its substantive results use one backbone family, roughly 270 optimizer steps, math tasks, and a particular LoRA training setup. The corruption band is deliberately enriched for transition-prone MATH problems, some likely present in pretraining, so its raw corruption counts are not a population forgetting rate. The authors also say their null does not address methods that generate new curricula.

It does not show that all self-improvement is illusory, that WTK has a noisy measurement path, or that a no-op control alone qualifies a package or target. External research is not a WTK Finding and does not promote the maturity of any WTK mechanism.

The builder impact

For WTK, keeping a baseline means evaluating the accepted, unchanged version through the same execution environment and schedule as the changed candidate. Recording both versions makes the comparison traceable. Repeated evaluation of the unchanged version estimates how much the metric varies without learning.

Our take: a candidate should not earn an improvement or promotion claim from a score change until the evaluation declares the no-op arm, metrics, seeds, budgets, scoring rules, and minimum effect required to support improvement in advance. A score difference alone is not enough if it cannot be distinguished from variation in the unchanged control.

The WTK test

WTK could run a bounded held-out evaluation with a frozen package, model binding, target projection, harness, evaluator, goal set, tool budget, and seed policy. Run the unchanged incumbent through every scheduled checkpoint alongside the changed candidate. Retain every attempt, score, receipt, and incomplete path, then report each claimed gain against its per-metric no-op distribution.

If the observed change does not exceed the predeclared threshold, the result is inconclusive, not evidence of improvement. That would not prove WTK's improvement loop is reliable. It would provide a measured comparison for the next improvement claim.

This complements testing harness changes on held-out tasks: task separation addresses evaluation on development examples, while an unchanged control measures score variation without improvement. Both inform the WTK research method, not automatic deployment approval.

Have an approach, result, or counterexample?

You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.

Contribute to this research question
Working with an AI assistant?

Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.

Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.

RECORD DETAILSReference RN-032
Artifact
Research Notes
Status
Published
Evidence posture
External research interpreted; proposed WTK experiment not yet run
Published
August 22, 2026
Author
WTK Research
Review
WTK human editorial review
Linked sources
1