WTKRESEARCH + ENGINEERING
← All engineering journal records

Research Notes

The Agents Did All the Engineering and None of the Science

In two case studies, frontier agents received six days and substantial compute, completed the engineering, and made no substantial progress on the research questions. A second model and scaffold reproduced the failure pattern.

Thursday's lead paper is the kind of result we make ourselves sit with, because it draws a boundary around one of WTK's favorite claims.

The catch: two cases where structure did not produce research progress

Researchers handed frontier agents the central open question from two unpublished NeurIPS submissions, six days of runtime, and thousands of dollars of compute. (Can AI agents conduct open-ended AI research?) The agents completed all the engineering unassisted. And they made no substantial progress on the actual research questions; both outputs were unambiguously rejected by the original authors. The detail that matters most to us: a robustness check with a second model and a second scaffold reproduced the failures. Changing the harness didn't move the outcome.

This paper does not test WTK and cannot establish a general class of goals where governance or composition adds no value. It does show two important counterexamples: checkable engineering completion did not imply open-ended research progress, and changing the model and scaffold did not remove the observed failure pattern.

How WTK research responds

By scoping the claim instead of defending it. We now distinguish contract-checkable goals, where an acceptance test can be written down and a structural-benefit claim can be tested, from judgment-terminal goals, where we currently lack an adequate test and will not claim one. WTK's floor-raising hypothesis remains testable for contract-checkable goals. It is not established for either goal class and is currently unsupported for judgment-terminal goals. Our quality numbers will name their goal class. A claim that states its own boundary is stronger, not weaker.

There's a second gift in the paper: five recurring failure modes, including ineffective backtracking from dead ends, poor resource awareness, and instruction drift. Three of those are repair moves WTK's bounded improvement loop doesn't have words for yet, so they went straight into the design pile for its repair vocabulary. Usual caveats: two case studies, one lab, so this is a boundary marker, not a ceiling proof.

Also on the radar

  • The adversary isn't standing still. GPT-Red describes a red-teaming agent trained via self-play that finds more successful attacks than human red-teamers and generalizes to held-out defender models and harnesses. A held-out adversarial test can report zero observed failures only against its named attack set and adversary tier; it is not a general safety percentage. WTK should therefore state blocking safety checks as zero tolerated failures on declared tests, with the adversary model attached. The never-averaged rule doesn't change; what the result is allowed to claim does. Caveats: it's OpenAI reporting on its own systems, and it's prompt-injection-specific.
  • A thousand-year-old provenance system. The Isnad-Rijal framework transfers classical hadith methodology to multi-agent knowledge: every claim carries a complete transmission chain, chains are judged at their weakest link, and compiled knowledge is routed three ways: serve, review, or quarantine. The chain half resembles WTK's intended grounding model; under a declared receipt contract, an incomplete chain can be rejected structurally. The genuinely new bit for us is that third routing state. Our grounding verdicts are binary today, and a middle state between serve and reject is a shape we're now considering. What we won't borrow: their per-narrator reliability scores. Grading trust and averaging it is exactly what WTK's validity rules exist to prevent.
  • When the rubric learns to go easy on you. DecoEvo names a failure mode we never want to meet: co-evolve your evaluation rubric with your solver, and apparent progress can come from the criteria getting easier. Their guard is that rubric updates are never driven by aggregate solver scores. We recorded that as a standing invariant before anyone here builds evaluator evolution: if judged criteria ever become evolvable in WTK, pass rates don't get a vote. The deterministic floor isn't a rubric and never becomes evolvable at all.

The takeaway

Good systems know what they can't do. This week WTK's claims got three clearer boundary markers: which goal classes WTK can currently test for structural benefit, which adversary a blocking safety test covered, and what a grounding verdict does and doesn't establish. None of that weakens the system. It's the difference between a claim and a brag.

See you at the next catch.

RECORD DETAILSReference RN-004
Artifact
Research Notes
Status
Published
Published
July 30, 2026
Linked sources
4