<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>WTK Engineering Journal</title><link>https://wtk-uplink.dev/blog/</link><description>Dated research notes, factory logs, field notes, findings, and failure reports from the WTK engineering workbench.</description><language>en-us</language><lastBuildDate>Sat, 01 Aug 2026 00:00:00 GMT</lastBuildDate><atom:link href="https://wtk-uplink.dev/rss.xml" rel="self" type="application/rss+xml"/><item><title>RN-005: One Poisoned Source Is All It Takes</title><link>https://wtk-uplink.dev/blog/one-poisoned-source/</link><guid isPermaLink="true">https://wtk-uplink.dev/blog/one-poisoned-source/</guid><pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate><category>Research Notes</category><description>Our daily sweep of new AI research turned up a number that should worry anyone building research agents: 54.7%.</description></item><item><title>RN-004: The Agents Did All the Engineering and None of the Science</title><link>https://wtk-uplink.dev/blog/all-engineering-no-science/</link><guid isPermaLink="true">https://wtk-uplink.dev/blog/all-engineering-no-science/</guid><pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate><category>Research Notes</category><description>In two case studies, frontier agents received six days and substantial compute, completed the engineering, and made no substantial progress on the research questions. A second model and scaffold reproduced the failure pattern.</description></item><item><title>RN-003: &quot;Held-Out&quot; Doesn&apos;t Mean Clean</title><link>https://wtk-uplink.dev/blog/held-out-doesnt-mean-clean/</link><guid isPermaLink="true">https://wtk-uplink.dev/blog/held-out-doesnt-mean-clean/</guid><pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate><category>Research Notes</category><description>In one multimodal fact-checking study, up to 29% of post-cutoff claims remained potentially contaminated, and the effect changed system rankings. That challenges how comparative value is measured, including in WTK.</description></item><item><title>RN-002: Verify Once, Answer Forever?</title><link>https://wtk-uplink.dev/blog/verify-once-answer-forever/</link><guid isPermaLink="true">https://wtk-uplink.dev/blog/verify-once-answer-forever/</guid><pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate><category>Research Notes</category><description>A frozen 12B model claims perfect accuracy at zero tokens by replaying verified answers. Big if true, and &quot;if&quot; is doing a lot of work.</description></item><item><title>RN-001: The Best Controls Might Be the Ones You Can&apos;t Ship</title><link>https://wtk-uplink.dev/blog/controls-that-cant-travel/</link><guid isPermaLink="true">https://wtk-uplink.dev/blog/controls-that-cant-travel/</guid><pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate><category>Research Notes</category><description>A new paper puts agent decision-making inside the model&apos;s hidden states. That&apos;s where WTK requires a tighter boundary, and we can now say why.</description></item></channel></rss>