WTKRESEARCH + ENGINEERING

Open Questions in Accountable AI Agent Engineering

Current questions and proposed experiments.

Active questions organize the experiments. Each names the next test, the current evidence boundary, and the result that would force the architecture to change.

MECHANISM IMPLEMENTED · CONTINUITY EXPERIMENT OWED

Can a governed goal remain measurable from intake through qualification and improvement?

Metrics can become proxies that reward test performance while degrading the real objective.

Why WTK investigates this
Goal-conditioned evaluation ties quality claims to the outcome the system was created to achieve.
Current implementation basis
WTK now compiles confirmed purpose into a target-neutral, versioned contract that preserves beneficiaries, responsibilities, success signals, counter-indicators, authority, cadence, and stop conditions. That contract is bound into packages, agent and team contracts, and goal-derived evaluation plans. Held-out evidence is sealed to exact goal, package, contract, plan, runtime, and evaluator identities, while malformed or invented purpose state fails closed.
Evidence boundary
WTK has not yet connected every purpose clause to derived work, graph nodes, handoffs, results, and deployment outcomes. It also has not demonstrated that conclusions remain comparable when the model, team shape, or execution harness changes while purpose remains constant.
What would change the approach
Repeated evidence that goal-derived evaluations fail to predict real outcome quality.
Next experiment
Hold the goal contract constant while changing model, team shape, and harness; measure whether conclusions remain traceable and useful.
BOUNDED MULTI-TARGET ASSURANCE · EQUIVALENCE GAP OPEN

Can an agent or team retain governed identity without belonging to a single runtime?

Harness-specific tools, memory, and orchestration semantics may be too important to abstract without loss.

Why WTK investigates this
Canonical packages and digests can separate source truth from runtime projection.
Current implementation basis
WTK now preserves exact package and contract identity across projection, target assurance, runtime evidence, control claims, and execution forms. Its catalog-release process uses preallocated proof attempts, retains failures, and evaluates seven readiness dimensions. Two exact-digest Local releases have passed that process, and one standalone package has bounded runtime-observed evidence on both Local and AgentCore.
Evidence boundary
These results show that WTK can detect, bind, and disclose target-specific differences. They do not establish general outcome equivalence. Hosted deployment, broader task and tool classes, parallel and fan-in team semantics, and uniform enforcement across runtimes remain open.
What would change the approach
Material governance properties cannot be represented or verified across supported targets.
Next experiment
Project the same package into distinct harnesses and compare authority, evidence, failure, and outcome semantics.
MECHANISM IMPLEMENTED · HUMAN STUDY OWED

Can discovery remain useful without allowing registration to imply readiness?

People routinely interpret catalog presence, ratings, and badges as institutional approval.

Why WTK investigates this
Explicit lifecycle and qualification fields make provenance and limitations inspectable.
Current implementation basis
WTK structurally separates registration, qualification, compatibility, promotion, and deployment readiness. Ordinary publication is distinct from a deployable catalog release. A deterministic four-case comparison also found that WTK's capability ladder preserved fit, cost consent, discovery, and construction boundaries in 4 of 4 declared decisions, while an installed-first shortcut made 0 of 4.
Evidence boundary
The comparison tested routing, not live fit judgment, setup, or accomplishment. No external non-expert study has established that ordinary users correctly distinguish ‘listed,’ ‘qualified,’ ‘works on my target,’ and ‘accomplishes my purpose.’ Comprehension and misuse resistance remain unproven.
What would change the approach
Users continue to treat available or registered components as fit and trusted despite the separation and warnings.
Next experiment
Run comprehension and live provisioning studies with operators who have not learned WTK terminology.
INTEGRITY MECHANISM IMPLEMENTED · RELIABILITY TRIAL OWED

When does independent review add reliability beyond receipt-grounded evaluation?

Panels may share blind spots, while evidence-starved judges can agree confidently on the same unsupported conclusion.

Why WTK investigates this
Receipt-backed grounding establishes what a judge can assess; independent bindings, quorum, and disagreement can then expose residual uncertainty.
Current implementation basis
WTK separates evidence by authority, seals held-out executions, and retains failed or inconclusive attempts. A 24-attempt causal probe added counted evidence: both bindings tracked supplied positive observations, but one failed all six strict absence cases while the comparison binding satisfied all twelve conditions. Deterministic and receipt failures remain separate from favorable semantic judgment.
Evidence boundary
The causal probe covered one task and two bindings. It did not establish source truth, general faithfulness, or the incremental value of panels. The decisive tribunal trial must still compare receipts versus no receipts across single judges, homogeneous panels, and diverse panels using reliability, disagreement, cost, and latency measures.
What would change the approach
After evidence supply is controlled, panels provide no useful reliability signal beyond a calibrated single judge, or their cost and latency outweigh any measured benefit.
Next experiment
Compare {receipts, no receipts} × {single judge, homogeneous panel, diverse panel} on false-pass, false-fail, abstention, calibration, disagreement, cost, and latency.
BOUNDED GOVERNED LOOP · LONGITUDINAL TEST OWED

Can agents, teams, and the factory improve without granting themselves authority to redefine success?

The evaluator, data, and promotion policy can drift together and normalize regressions.

Why WTK investigates this
Versioned candidates, immutable evidence, qualification, and explicit promotion keep learning reviewable.
Current implementation basis
WTK separates incumbent and candidate versions, keeps feedback advisory, and preserves operator-controlled promotion. Two deterministic experiments now strengthen that mechanism: typed repair routing selected the declared layer in 5 of 5 persistent-failure cases without repeating instruction steering, and a package edit preserved prior artifacts while rejecting their construction and approval authority for the revised digest.
Evidence boundary
The routing experiment did not apply a repair, and the retirement experiment did not complete a full rebuild and readmission. Repeated improvement across many generations, evaluator drift, governance recursion, adversarial pressure, accumulated bad changes, and policy-based autonomous promotion remain untested.
What would change the approach
The loop cannot reliably detect regressions, retire stale evidence, or preserve a recoverable qualified incumbent.
Next experiment
Apply typed repairs to retained live failures, rebuild a revised qualified package, then run long-duration improvement cycles with drift, rollback, and an unchanged incumbent.
One additional research question is in preparation
CONTROLS IMPLEMENTED · ADVERSARIAL SYSTEM TEST OWED

Can a team become unsafe even when every individual action appears policy-compliant?

Current implementation basis
WTK implements delegation authority, budgets, dependency blocking, lineage, typed handoffs, and fan-in controls.
Evidence boundary
No collusive or emergent-harm experiment has tested a policy-compliant team as one adversarial system.
What would change the approach
System-level harmful behavior remains invisible to the proposed evidence and enforcement boundaries.
Next experiment
Use collusive, incremental, and shared-state adversarial scenarios against the team as a system.
Research contribution

Conditions for further automation

These questions extend the existing challenge program. They do not create a second ledger or imply that policy-authorized promotion is ready.

What evidence permits a factory function to move from human approval to human oversight?

Human trust calibration · Yardstick validity

Which decisions should never receive unattended promotion authority?

Governance independence · Organizational acceptance

Can an automated promotion plane remain independent from the factory that supplies its candidates?

Governance independence · Authority concentration

How should autonomy be reduced after drift, uncertainty, evidence loss, or policy change?

Long-term drift · Evidence authenticity

How do evaluator drift and test saturation affect long-running improvement loops?

Yardstick validity · Repair-loop integrity

Does progressive automation create enough net value to justify its governance cost?

Economic value

Can automated remediation remain bounded without hiding repeated failure?

Repair-loop integrity · Falsifiability in practice
Review the governed self-improvement model