WTKRESEARCH + ENGINEERING

ACTIVE RESEARCH

The roadmap is what we still do not know.

Five questions currently organize the experiments. Each names the next test and the evidence that would force the architecture to change.

ACTIVE · PARTIAL EVIDENCE

Can a governed goal remain measurable from intake through qualification and improvement?

Current implementation basis
Goal-derived evaluation plans execute against captured sessions and feed package qualification in the current WTK implementation.
Evidence boundary
This demonstrates a bounded implementation path, not continuity across every harness, model, or long-running improvement cycle.
Next experiment
Hold the goal contract constant while changing model, team shape, and harness; measure whether conclusions remain traceable and useful.
ACTIVE · OPEN EQUIVALENCE GAP

Can an agent or team retain governed identity without belonging to a single runtime?

Current implementation basis
Current WTK projections preserve package identity and contract hashes across a bounded set of target adapters.
Evidence boundary
Broader cross-target outcome equivalence and uniform enforcement are explicitly open.
Next experiment
Project the same package into distinct harnesses and compare authority, evidence, failure, and outcome semantics.
MECHANISM IMPLEMENTED · HUMAN STUDY OWED

Can discovery remain useful without allowing registration to imply readiness?

Current implementation basis
WTK structurally separates registration, qualification, compatibility, and promotion in its current registry model.
Evidence boundary
No comprehension or misuse study has established that ordinary operators understand and preserve that distinction.
Next experiment
Run comprehension and misuse studies with operators who have not learned WTK terminology.
ACTIVE · RELIABILITY TRIALS OWED

Do independent judges reduce semantic evaluation error without creating majority-vote theater?

Current implementation basis
WTK implements separate review lenses, judge quorum, disagreement reporting, and a deterministic floor that semantic judgment cannot erase.
Evidence boundary
No controlled study yet shows that a diverse panel is more reliable than a calibrated single judge.
Next experiment
Compare single-judge, homogeneous-panel, and diverse-panel reliability against held-out human-grounded cases.
BOUNDED EVIDENCE · LONGITUDINAL TEST OWED

Can agents, teams, and the factory improve without granting themselves authority to redefine success?

Current implementation basis
WTK implements incumbent and candidate separation, regression comparison, rollback, qualification, and operator-controlled promotion, with bounded demonstrations.
Evidence boundary
Long-duration drift, governance recursion, and repeated self-improvement safety have not been established.
Next experiment
Run long-duration improvement cycles with adversarial goals, evaluator drift, rollback, and incumbent comparison.
One additional research question is in preparation
CONTROLS IMPLEMENTED · ADVERSARIAL SYSTEM TEST OWED

Can a team become unsafe even when every individual action appears policy-compliant?

Current implementation basis
WTK implements delegation authority, budgets, dependency blocking, lineage, typed handoffs, and fan-in controls.
Evidence boundary
No collusive or emergent-harm experiment has tested a policy-compliant team as one adversarial system.
Next experiment
Use collusive, incremental, and shared-state adversarial scenarios against the team as a system.