Measuring Agent Research Progress Beyond Activity Counts
More agent work is useful only when we can tell what it accomplished and what people had to correct.
A factory can generate more code and run more experiments without producing more trustworthy results. To understand whether agents improve the work, we need to inspect what succeeded and how much correction it required.
For WTK, this becomes more important as automation increases. Counting completed sessions is convenient. Establishing that those sessions produced usable, permitted outcomes takes more evidence.
What OpenAI measured
OpenAI's September 6 report, Research acceleration: The view inside OpenAI, distinguishes growing internal agent activity from harder-to-measure research progress.
Its task analysis samples 2% of January-to-July conversations, selects interactive coding-agent sessions, and excludes automations. An AI-based evaluator checks sessions and subsequent evidence. OpenAI reports agreement with human judgments on 25 manually labeled tasks: limited validation, not a broad reliability guarantee. The main task-outcome plots exclude uncertain outcomes. Duration means estimated human work, not agent runtime.
OpenAI reports that over half of successful tasks in the estimated four-to-eight-hour category involved at least one human intervention. This measures the occurrence of intervention, not the time spent correcting work or total human effort. It does not establish whether the assisted workflow saved time overall.
Tools, usage, task selection, and compute changed during this observational period. These measurements do not establish a causal productivity gain for WTK or evaluate an unattended factory.
Count the correction, not just the first answer
Consider a hypothetical research agent that finishes a comparison in ten minutes. Its reviewer then spends an hour replacing unsupported citations and rerunning the analysis. Another configuration takes longer initially but produces a result requiring little correction.
The first completion timestamp cannot settle which approach is better. Neither can a correction count: one intervention might take a minute or an hour. The time accounting in this example is our proposed measurement distinction, not a result established by OpenAI's intervention figures.
Our runtime evaluation note asks whether a supported configuration can do the intended work. This adds a measurement question: does the review follow the result far enough to see corrections and unresolved outcomes?
For WTK research, we would keep the intended goal, the acceptance evidence, the agent's resource use, and the operator's corrective work distinguishable. A successful task after intervention can still be useful. It should not be reported as unattended success.
Human approval is different from corrective work. Reviewing a proposed publication or authorizing a deployment may be a deliberate control, not a failure of automation. Counting all human involvement as an error would reward bypassing the very governance we want to preserve.
Keep unresolved outcomes visible
An outcome we cannot verify belongs in an unresolved category. Excluding it from an analysis may answer a narrower question, but the excluded share and its possible effect should remain visible.
Our research method calls for bounded claims and retained attempts. That principle also applies to productivity: later success should not erase the failed attempts or human repair that preceded it.
This note does not establish a WTK productivity gain, recommend removing reviewers, or authorize a research run. Before proposing a comparison, we need a practical way to record correction effort and follow-up outcomes without collecting unnecessary private information.
The takeaway is straightforward: measure useful results, report the work needed to obtain them, and preserve approval boundaries. More activity can be a helpful signal, but it is not the outcome we are trying to improve.
Source
OpenAI, Research acceleration: The view inside OpenAI, September 6, 2026, including its task-outcome methods disclosure. The factory example and WTK measurement questions are our interpretation, not findings reported for WTK.
Have an approach, result, or counterexample?
You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.
Contribute to this research question →Working with an AI assistant?
Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.
Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.