One Poisoned Source Is All It Takes
Our daily sweep of new AI research turned up a number that should worry anyone building research agents: 54.7%.
Open record →RESEARCH · EXPERIMENTS · FINDINGS
WTK tests claims about autonomous agent systems against credible research, opposing arguments, and experiments capable of changing the architecture.
RECENT RECORDS
Each note states what the evidence permits, where it stops, and what builders should test next.
Our daily sweep of new AI research turned up a number that should worry anyone building research agents: 54.7%.
Open record →In two case studies, frontier agents received six days and substantial compute, completed the engineering, and made no substantial progress on the research questions. A second model and scaffold reproduced the failure pattern.
Open record →In one multimodal fact-checking study, up to 29% of post-cutoff claims remained potentially contaminated, and the effect changed system rankings. That challenges how comparative value is measured, including in WTK.
Open record →ACTIVE EXPERIMENTS
The active questions name the next experiment and the evidence that would force an architectural change.
Next experimentHold the goal contract constant while changing model, team shape, and harness; measure whether conclusions remain traceable and useful.
Next experimentProject the same package into distinct harnesses and compare authority, evidence, failure, and outcome semantics.
Next experimentRun comprehension and misuse studies with operators who have not learned WTK terminology.
REPEATABLE METHOD
Signal, evidence, boundary, builder impact, and the next WTK test remain visible together.
ARCHITECTURE
The factory, governance model, and reference patterns remain available when a reader wants the implementation detail. They no longer interrupt the research path.