One Poisoned Source Is All It Takes
Our daily sweep of new AI research turned up a number that should worry anyone building research agents: 54.7%.
Read the record →WTK ENGINEERING JOURNAL
Notes, findings, failures, and factory logs remain one chronological record of what WTK is learning and changing.
The most useful catches challenge a WTK assumption, narrow a claim, or define the next experiment.
Our daily sweep of new AI research turned up a number that should worry anyone building research agents: 54.7%.
Read the record →In two case studies, frontier agents received six days and substantial compute, completed the engineering, and made no substantial progress on the research questions. A second model and scaffold reproduced the failure pattern.
Read the record →In one multimodal fact-checking study, up to 29% of post-cutoff claims remained potentially contaminated, and the effect changed system rankings. That challenges how comparative value is measured, including in WTK.
Read the record →A frozen 12B model claims perfect accuracy at zero tokens by replaying verified answers. Big if true, and "if" is doing a lot of work.
Read the record →A new paper puts agent decision-making inside the model's hidden states. That's where WTK requires a tighter boundary, and we can now say why.
Read the record →