The roadmap is what we still do not know.
Five questions currently organize the experiments. Each names the next test and the evidence that would force the architecture to change.
ACTIVE · PARTIAL EVIDENCE
Can a governed goal remain measurable from intake through qualification and improvement?
- Current implementation basis
- Goal-derived evaluation plans execute against captured sessions and feed package qualification in the current WTK implementation.
- Evidence boundary
- This demonstrates a bounded implementation path, not continuity across every harness, model, or long-running improvement cycle.
- Next experiment
- Hold the goal contract constant while changing model, team shape, and harness; measure whether conclusions remain traceable and useful.
ACTIVE · OPEN EQUIVALENCE GAP
Can an agent or team retain governed identity without belonging to a single runtime?
- Current implementation basis
- Current WTK projections preserve package identity and contract hashes across a bounded set of target adapters.
- Evidence boundary
- Broader cross-target outcome equivalence and uniform enforcement are explicitly open.
- Next experiment
- Project the same package into distinct harnesses and compare authority, evidence, failure, and outcome semantics.
MECHANISM IMPLEMENTED · HUMAN STUDY OWED
Can discovery remain useful without allowing registration to imply readiness?
- Current implementation basis
- WTK structurally separates registration, qualification, compatibility, and promotion in its current registry model.
- Evidence boundary
- No comprehension or misuse study has established that ordinary operators understand and preserve that distinction.
- Next experiment
- Run comprehension and misuse studies with operators who have not learned WTK terminology.
ACTIVE · RELIABILITY TRIALS OWED
Do independent judges reduce semantic evaluation error without creating majority-vote theater?
- Current implementation basis
- WTK implements separate review lenses, judge quorum, disagreement reporting, and a deterministic floor that semantic judgment cannot erase.
- Evidence boundary
- No controlled study yet shows that a diverse panel is more reliable than a calibrated single judge.
- Next experiment
- Compare single-judge, homogeneous-panel, and diverse-panel reliability against held-out human-grounded cases.
BOUNDED EVIDENCE · LONGITUDINAL TEST OWED
Can agents, teams, and the factory improve without granting themselves authority to redefine success?
- Current implementation basis
- WTK implements incumbent and candidate separation, regression comparison, rollback, qualification, and operator-controlled promotion, with bounded demonstrations.
- Evidence boundary
- Long-duration drift, governance recursion, and repeated self-improvement safety have not been established.
- Next experiment
- Run long-duration improvement cycles with adversarial goals, evaluator drift, rollback, and incumbent comparison.
One additional research question is in preparation
CONTROLS IMPLEMENTED · ADVERSARIAL SYSTEM TEST OWEDCan a team become unsafe even when every individual action appears policy-compliant?
- Current implementation basis
- WTK implements delegation authority, budgets, dependency blocking, lineage, typed handoffs, and fan-in controls.
- Evidence boundary
- No collusive or emergent-harm experiment has tested a policy-compliant team as one adversarial system.
- Next experiment
- Use collusive, incremental, and shared-state adversarial scenarios against the team as a system.