A supplied-text-only team acquired unnecessary capability requirements during planning, before its agents could execute.
Next experiment
Compare current WTK capability selection with explicit negative-constraint preservation across supplied and live inputs, single agents and teams, and held-out paraphrases.
Later user trials exposed a completion path that needed to distinguish a revised goal from its earlier successful package.
Next experiment
Compare current WTK completion handling with persisted intent-bound completion checks across unchanged, materially revised, and resumed single-agent and team goals.
A proposal to account for every required team role before assessing the executed workflow.
Next experiment
Retain one complete control and matched broken-lane variants with sanitized evidence identities, then verify that each incomplete condition blocks the deployable claim without blocking the control.
A repair can respond to a failed review without copying the withheld case that exposed it.
Next experiment
Run separately authorized fresh reviews of repairs derived from public regressions and measure whether they preserve the original boundary without exposing withheld inputs.
An improvement score is only meaningful when its intended outcome is declared before the comparison begins.
Next experiment
Run a separately authorized comparison that keeps one declared improvement definition, package pair, evaluation method, and complete attempt set fixed.
A second look can test retained evidence without quietly replacing the decision it was meant to examine.
Next experiment
Review an advisory-only replay path for rejected windows without changing their verdicts; any implementation or live replay requires separate approval.
Reviewing durable package source is not enough when the executable form also depends on compiled inputs.
Next experiment
Vary source, compiler input, target, and approval conditions independently while verifying that each catalog claim remains bound to the executable subject.
When a failure survives an instruction change, the system should change repair layers instead of polishing the same guess again.
Next experiment
Run counted repair sequences across instruction, output-contract, capability, evidence-policy, and goal failures while measuring escalation accuracy, regression, abstention, cost, and operator intervention.
Exploration, package construction, testing, and repair should remain one traceable journey without turning the conversation into source truth.
Next experiment
Resume the same governed work across interruption, revision, failed review, repair, and retest while checking that identity and authority remain attached to the correct artifacts.
Changing governed package source should create a new accountable revision, not let yesterday's proof follow it forward.
Next experiment
Apply controlled package changes across supported execution forms and verify that every affected test, qualification, projection, and catalog claim is invalidated and re-earned.
A registry should not promise an action that no compiled target can actually perform.
Next experiment
Hold one governed capability constant across two supported targets and compare executable binding, runtime receipts, refusal behavior, and target-specific limitations.
A tool connection is not ready merely because it was once configured or once passed a check.
Next experiment
Vary missing credentials, stale checks, revoked authority, failed retests, and usable connections while confirming that the original evidence obligation remains visible.
For a governed tool requirement, an answer is only a partial result until the runtime shows the tool was actually used.
Next experiment
Exercise the same governed tool requirement across additional tool shapes, denied actions, failures, and state-changing boundaries while retaining complete target-local receipts.
A missing evidence capability should trigger an explicit decision, not a quieter unsupported answer.
Next experiment
Vary unavailable, incompatible, unconfigured, and costly evidence capabilities while checking that the goal obligation remains visible and no unsupported answer path appears.
Changing the work should create a new accountable version, not let old evidence answer a new question.
Next experiment
Change one material goal condition at a time across multiple execution forms and verify that the new result remains traceable to its exact goal version without inheriting stale evidence.
A deployment claim is only as honest as the complete evidence set it keeps.
Next experiment
Change one declared runtime condition at a time and verify that the old evidence remains historical while the new claim stays blocked until its full planned set is complete.
The August 5 development record separates session-bound feedback from evaluation results; safe autonomous promotion remains unproven.
Next experiment
Test a preauthorized improvement envelope in advisory, shadow, and bounded autonomous modes while preserving evidence, incumbent comparison, rollback, and escalation.
WTK separated evidence by purpose so a convincing run cannot quietly become a qualification claim.
Next experiment
Hold the governed goal constant while changing execution conditions and verify that evidence authority, failed attempts, and qualification scope remain intact.