WTKRESEARCH + ENGINEERING
← Back to Journal

A Final Score Cannot Tell You What the Agent Learned

A retained lesson should have to beat a matched run that starts fresh.

A final score can tell us whether a run ended well. By itself, it cannot tell us whether the experience carried forward from that run improved the next decision, or merely made the next failure harder to notice.

The signal

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development studies seven frontier models on 36 long-horizon research-and-development tasks. Instead of treating the final outcome as the whole story, the authors use rule-based measures for solution framing, execution, and feedback control. They also compare matched conditions with and without retained experience within a task and with and without extracted lessons across tasks.

Their result is a useful correction to the default story about learning agents. Performance varied substantially across runs. Strong solutions mostly adapted or combined established techniques rather than producing methodological novelty. Under the paper's defined conditions, retained experience often improved a later decision and sometimes misled it. The effects varied by model, task, and representation of the retained material, so the results do not support a universal claim about memory or improvement.

The evidence

The paper's practical contribution is a measurement design, not only a measurement question. For within-task reuse, it branches from the same intermediate solution and compares the next commit with retained experience against a reinitialized condition that removes prior context, notes, and comments. For cross-task reuse, it holds the target task conditions fixed and compares runs with and without extracted lessons. A final score without that control can hide a good lesson, a bad carryover, or a lucky run.

For WTK, that matters because feedback and retained evidence should remain bounded inputs to a future candidate, not invisible authority to repeat a prior choice. Our take: a learning loop should name what crosses from one task to the next, preserve the source and scope of that material, and compare it with a declared no-reuse condition. Otherwise, calling later behavior an improvement is too easy.

The boundary

This is a preprint about the authors' selected models, tasks, harnesses, and experience conditions. Its controlled comparisons support model- and task-specific effects of the experience treatments the authors tested. They are not evidence that WTK's improvement loop is effective, that retained experience is safe or harmful in general, or that a WTK package should carry information across goals. The paper does not evaluate WTK package contracts, authority boundaries, target projections, independent qualification, or operator approval.

The paper does not isolate which individual lesson, note, or retained artifact caused a later result, nor does it show that its effect transfers to WTK's governed setting. Its comparisons hold the stated target conditions fixed, but they do not turn any particular reuse mechanism into a universal rule. It is therefore not a WTK Finding and does not promote the maturity of any WTK mechanism.

The builder impact

Builders can make reuse inspectable before they call it learning. Declare whether the next task receives a prior plan, feedback, tool result, summary, or candidate artifact. Bind that material to its source task and authority. Keep the no-reuse control, every attempted task, and every regression visible alongside the final outcomes.

That structure gives a future review something useful to ask: under these conditions, did this particular retained artifact improve the next bounded decision? It also keeps a promising prior answer from quietly rewriting the success criterion for the next one.

The WTK test

WTK could test the question with a small held-out comparison. Freeze a package, model binding, target harness, goals, evaluator plan, and budgets. Run one arm with a predeclared, provenance-bound retained artifact and one matched arm without it. Preserve each framing decision, execution receipt, feedback artifact, final result, and regression.

The result would support only a bounded conclusion about that artifact, task set, and execution form. Until that comparison exists, the paper is a reason to measure retained experience, not evidence that WTK has earned an improvement claim.

Have an approach, result, or counterexample?

You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.

Contribute to this research question
Working with an AI assistant?

Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.

Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.

RECORD DETAILSReference RN-026
Artifact
Research Notes
Status
Published
Evidence posture
Published with the evidence boundary stated in this record
Published
August 17, 2026
Author
WTK Research
Review
WTK human editorial review
Linked sources
1