WTKRESEARCH + ENGINEERING
← Back to Journal

A Replay Cannot Rewrite the First Verdict

A second look can test retained evidence without quietly replacing the decision it was meant to examine.

Review update: September 5, 2026

The original account below reports one local comparison and an evaluator-only replay. Reviewing the later August 24 window changes the follow-up: that window rejected the candidate despite a positive average score. The proposed favorable three-window result was not established. The original account and its then-proposed next hypothesis remain below as history, not the current work queue.

Reviewing again is different from running again

Suppose two retained agent outputs receive an initial comparison. A later evaluator sees the same outputs with their order changed and their version labels hidden. Agreement would concern those retained outputs under the two review conditions, not two fresh executions of the agents.

That distinction makes the replay useful for examining sensitivity to review conditions. It does not establish statistical independence, eliminate shared evaluator errors, or show that the agent improved on new work. A separately authorized reviewer can still share biases with the first.

The later rejection matters more than the average

The retained second-window report describes ten subject calls and three evaluator calls, all reaching terminal records, with no retries. Its candidate average was 0.8320 against 0.6660 for the incumbent. But the task requiring an operator decision when tool evidence was ambiguous regressed by 0.6667, beyond the predeclared 0.20 permitted regression. The candidate was rejected. The positive average could not override that task-level boundary.

Under a rule requiring every window to pass, a later good result cannot repair this rejection. This is a reported negative development result, not evidence that every candidate or improvement method fails.

The same report records an incomplete assurance step: the replay adapter accepted only passing source windows, so its zero-call check refused this rejected window. That conflicted with the protocol's requirement to replay every window. The source verdict was preserved, and neither the code nor rule was changed after seeing the result. Any correction and live replay need separate review and authorization; a replay must remain advisory even for a rejected source.

What is and is not available

The source reports name a fixed factory-coordinator version pair, a gpt-5.6-luna subject binding, and a gpt-5.6-terra evaluator binding. The first replay reused retained outputs and the same evaluator model with reordered presentation and hidden version labels. It was not an independent model or a new subject trial.

This editorial review inspected the retained reports and the answer-redacted second-window summary, whose digest matched the report. It did not open encrypted raw responses or make provider calls. The public article does not expose the complete original evidence bundle, so readers cannot independently reproduce the comparison from it. These August observations do not characterize current model performance.

The improvement-definition log explains why the objective must be declared before interpreting a score. Our research method requires retained unfavorable and inconclusive outcomes too.

A useful evidence supplement would expose sanitized paired grades, presentation order, evaluator settings, and all attempt dispositions. Where confidentiality prevents publishing subject content, the report must explain what cannot be independently checked. The next decision is whether to make rejected windows reviewable without changing their authority, not whether to replace an unfavorable window with another attempt.

Original account: August 24, 2026

Current state

A version comparison can produce a tempting score before anyone has shown that a later review would see the same thing. A rerun is not an independent look if it changes the task, the evidence, or the authority of the original decision.

Changes

One completed local engineering cycle kept the original comparison fixed, retained its exact subject outputs behind a restricted review boundary, and added a separately authorized evaluator-only replay. The replay used those retained outputs, changed the presentation order, hid which version produced each side, and made no new subject run.

The original verdict stayed immutable. The replay was appended as advisory evidence, so it could support, complicate, or disagree with the first direction without turning itself into candidate admission, qualification, registry promotion, or deployment authority.

Failures observed

The cycle leaves a useful limit visible. A single later comparison can agree overall while differing on individual tasks. That does not settle whether the score reflects lasting quality, a narrow task set, or a reviewer-specific tendency.

Assumptions removed

We cannot treat a score delta as self-explanatory, or treat a second review as permission to replace the first. Independent checking needs its own frozen method, retained evidence, and authority boundary.

Evidence

The completed cycle preallocated terminal evidence records before its planned work, completed the fixed comparison, then completed one evaluator-only replay over retained outputs. The replay did not rerun the subject and did not change package, qualification, registry, publication, or deployment state.

Limitations

This is one local engineering observation, not a reliability result. It does not show that independent replays are sufficient, that the review method is unbiased, that the score reflects production outcomes, or that the pattern transfers to other agents, tasks, reviewers, or targets.

Next hypothesis

Repeat the same frozen comparison in two later calendar windows and require one bounded independent replay per window. Keep agreements, disagreements, failed calls, and hard blocks visible, then report the narrow result without letting any replay rewrite the original verdict.

Have an approach, result, or counterexample?

You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.

Contribute to this research question
Working with an AI assistant?

Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.

Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.

RECORD DETAILSReference FL-020
Artifact
Factory Logs
Status
Published
Evidence posture
Historical engineering account updated with a later rejection and incomplete replay requirement
Published
August 24, 2026
Author
WTK Research
Review
WTK human editorial review
Linked sources
None declared