WTKRESEARCH + ENGINEERING

Autonomous Agent Factory Engineering Logs

The factory, under construction.

Engineering records document what changed, what failed, which assumptions were removed, what evidence was produced, and what the next build must test.

Factory Logs

No Tools Has to Survive the Factory

A supplied-text-only team acquired unnecessary capability requirements during planning, before its agents could execute.

Next experiment
Compare current WTK capability selection with explicit negative-constraint preservation across supplied and live inputs, single agents and teams, and held-out paraphrases.
Full article
Factory Logs

A Revised Goal Cannot Inherit Yesterday's Success

Later user trials exposed a completion path that needed to distinguish a revised goal from its earlier successful package.

Next experiment
Compare current WTK completion handling with persisted intent-bound completion checks across unchanged, materially revised, and resumed single-agent and team goals.
Full article
Factory Logs

A Retry Should Not Erase the First Attempt

Hidden transport retries can separate the work an agent performs from the attempts its factory records.

Next experiment
Compare current WTK with explicit cross-layer attempt reconciliation under controlled faults, retaining useful-task controls and every attempt.
Full article
Factory Logs

Proposed Evidence Checks for Agent Teams

A proposal to account for every required team role before assessing the executed workflow.

Next experiment
Retain one complete control and matched broken-lane variants with sanitized evidence identities, then verify that each incomplete condition blocks the deployable claim without blocking the control.
Full article
Factory Logs

A Repair Cannot Borrow Its Test

A repair can respond to a failed review without copying the withheld case that exposed it.

Next experiment
Run separately authorized fresh reviews of repairs derived from public regressions and measure whether they preserve the original boundary without exposing withheld inputs.
Full article
Factory Logs

A Better Claim Needs Its Own Definition

An improvement score is only meaningful when its intended outcome is declared before the comparison begins.

Next experiment
Run a separately authorized comparison that keeps one declared improvement definition, package pair, evaluation method, and complete attempt set fixed.
Full article
Factory Logs

A Replay Cannot Rewrite the First Verdict

A second look can test retained evidence without quietly replacing the decision it was meant to examine.

Next experiment
Review an advisory-only replay path for rejected windows without changing their verdicts; any implementation or live replay requires separate approval.
Full article
Factory Logs

A Valid Error Is Not a Passing Answer

A response can match a schema and still fail the obligation that the fixture was meant to test.

Next experiment
Replay paired ordinary-success and expected-rejection fixtures across package shapes and targets without allowing either result to satisfy the other.
Full article
Factory Logs

A Secret Is Not a Tool Setting

A connection can name the credential it needs without becoming the place that governs the credential itself.

Next experiment
Exercise the same credential boundary across target setup, revocation, and recovery paths while retaining each authorization decision.
Full article
Factory Logs

A Catalog Claim Needs the Source It Will Run

Reviewing durable package source is not enough when the executable form also depends on compiled inputs.

Next experiment
Vary source, compiler input, target, and approval conditions independently while verifying that each catalog claim remains bound to the executable subject.
Full article
Factory Logs

A File Tool Has to Stop at the Boundary

Read-only describes an operation. A package grant must also say where that operation may reach.

Next experiment
Compare the same package grant across distinct target harnesses and retain every authorization, refusal, and receipt.
Full article
Factory Logs

A Registry Miss Can Be a Language Miss

A complete catalog still fails when a user's words never reach the capability it already has.

Next experiment
Expand the fixed probe set with predeclared paraphrases and measure precision before changing consumer-language detection.
Full article
Factory Logs

A Gate Failure Has to Name Its Own Layer

A blocked candidate is not automatically evidence that the candidate is the problem.

Next experiment
Use fixed duplicate-name and paired fail-closed cases to verify that generation and scoring failures are attributed to their own layers.
Full article
Factory Logs

The Repair Loop Needs an Exit Ramp

When a failure survives an instruction change, the system should change repair layers instead of polishing the same guess again.

Next experiment
Run counted repair sequences across instruction, output-contract, capability, evidence-policy, and goal failures while measuring escalation accuracy, regression, abstention, cost, and operator intervention.
Full article
Factory Logs

One Work Thread Should Not Become Two Histories

Exploration, package construction, testing, and repair should remain one traceable journey without turning the conversation into source truth.

Next experiment
Resume the same governed work across interruption, revision, failed review, repair, and retest while checking that identity and authority remain attached to the correct artifacts.
Full article
Factory Logs

An Edit Has to Retire Its Old Evidence

Changing governed package source should create a new accountable revision, not let yesterday's proof follow it forward.

Next experiment
Apply controlled package changes across supported execution forms and verify that every affected test, qualification, projection, and catalog claim is invalidated and re-earned.
Full article
Factory Logs

A Capability Claim Has to Reach an Executable Path

A registry should not promise an action that no compiled target can actually perform.

Next experiment
Hold one governed capability constant across two supported targets and compare executable binding, runtime receipts, refusal behavior, and target-specific limitations.
Full article
Factory Logs

A Green Check Can Go Stale

A tool connection is not ready merely because it was once configured or once passed a check.

Next experiment
Vary missing credentials, stale checks, revoked authority, failed retests, and usable connections while confirming that the original evidence obligation remains visible.
Full article
Factory Logs

An Answer Can Be Correct and Still Miss the Required Tool

For a governed tool requirement, an answer is only a partial result until the runtime shows the tool was actually used.

Next experiment
Exercise the same governed tool requirement across additional tool shapes, denied actions, failures, and state-changing boundaries while retaining complete target-local receipts.
Full article
Factory Logs

When a Required Tool Is Missing, Stop Guessing

A missing evidence capability should trigger an explicit decision, not a quieter unsupported answer.

Next experiment
Vary unavailable, incompatible, unconfigured, and costly evidence capabilities while checking that the goal obligation remains visible and no unsupported answer path appears.
Full article
Factory Logs

A Revised Goal Must Keep Its History

Changing the work should create a new accountable version, not let old evidence answer a new question.

Next experiment
Change one material goal condition at a time across multiple execution forms and verify that the new result remains traceable to its exact goal version without inheriting stale evidence.
Full article
Factory Logs

A Release Claim Has to Remember the Misses

A deployment claim is only as honest as the complete evidence set it keeps.

Next experiment
Change one declared runtime condition at a time and verify that the old evidence remains historical while the new claim stays blocked until its full planned set is complete.
Full article
Factory Logs

No Evidence Path, No Confident Answer

WTK now turns a missing required evidence path into a recorded limitation instead of accepting unsupported confident prose.

Next experiment
Vary missing, empty, failed, irrelevant, and path-mismatched evidence across agents, teams, and harnesses.
Full article
Factory Logs

Feedback Can Guide Improvement Without Rewriting the Verdict

The August 5 development record separates session-bound feedback from evaluation results; safe autonomous promotion remains unproven.

Next experiment
Test a preauthorized improvement envelope in advisory, shadow, and bounded autonomous modes while preserving evidence, incumbent comparison, rollback, and escalation.
Full article
Factory Logs

A Package Build Is Not a Deployment Claim

WTK now requires evidence from the compiled execution form before a catalog entry can claim deployable readiness.

Next experiment
Apply the same release boundary to additional harnesses, team topologies, tools, and operating conditions.
Full article
Factory Logs

Making Evaluation Evidence Harder to Game

WTK separated evidence by purpose so a convincing run cannot quietly become a qualification claim.

Next experiment
Hold the governed goal constant while changing execution conditions and verify that evidence authority, failed attempts, and qualification scope remain intact.
Full article