WTKRESEARCH + ENGINEERING

Challenges to Governed Autonomous Agent Factories

Try to prove it wrong.

Start with the largest unresolved challenges to the architecture. Open a record only when you want its evidence, limits, and related questions.

21 shown

THE TEN

Largest unresolved challenges.

Ranked by how strongly each question could change the architecture. Open a record for the reasoning and evidence boundary.

RANK 1
STATE · namedMOVES THROUGH · operational

Falsifiability in Practice

What finding would force a change to the architecture, rather than a narrower claim?

Why it exists and what moves it +

Recent review cycles produced disconfirming evidence and narrower boundary statements, but no precommitted condition determined when the architecture itself must change. Narrowing a claim is legitimate work, but a program that can always shrink its promise cannot expose architectural failure.

This challenge remains first until a trip-wire is written in advance that names a class of result capable of changing the architecture itself.

Share an approach or counterexample
RANK 2
STATE · namedMOVES THROUGH · operational

Yardstick Validity

Can the measurements that grade these challenges be trusted?

Why it exists and what moves it +

Every state in this program is downstream of evaluation. Held-out sets do not yet have to disclose contamination controls, and the stability of a verifier's repeated judgments has not been audited.

On 2026-08-14, a 24-attempt causal probe showed that both tested model bindings tracked supplied positive observations, while one binding failed all six strict evidence-absence cases. This demonstrates that positive receipt use did not ensure reliable abstention or receipt-reference behavior in the tested task. It does not establish evaluator stability or the incremental value of a review panel.

This is separate from authority over the qualifier. A legitimate authority can still hold a crooked ruler. See C-02.

Share an approach or counterexample
RANK 3
STATE · boundedMOVES THROUGH · research

Input Truth

Can a system with structurally perfect grounding still produce confidently wrong work?

Why it exists and what moves it +

A 2026 study introduced one credible but misleading document into deep-research sources and raised mean report-level false-conclusion adoption from 0% to 54.7% across the studied configurations. Source authority and presentation mattered, while search rank and additional documents had limited influence.

WTK is designed to provide structured traceability to the sources recorded for a bounded execution; it does not establish source truth. Cross-model verification alone did not close the studied gap, but this result does not establish that every independent evaluation design will fail. The study covers selected agents, backbones, and constructed misleading documents.

Narrowing relationshipThis challenge partially narrows C-05.

Share an approach or counterexample
RANK 4
STATE · namedMOVES THROUGH · operational

Governance Independence

Who qualifies the qualifier, evaluates the tribunal, and promotes changes to governance?

Why it exists and what moves it +

An infinitely recursive governance hierarchy is impossible. The architecture must terminate in an explicit trust root while preventing the governed system from silently moving that root.

Falsification signal: A governed artifact can modify its evaluator, fixtures, promotion criteria, evidence authority, or policy root and then use the modified system to authorize itself.

Evidence direction: Define the terminal trust root, separation-of-duty matrix, governance-change process, external anchors, emergency authority, and the evidence required to change any of them.

Authority over the qualifier and validity of the measurement are separate failures. See C-18.

Share an approach or counterexample
RANK 5
STATE · boundedMOVES THROUGH · research

Goal-Class Boundedness

Is there a class of goals where governance adds cost and no floor?

Why it exists and what moves it +

A 2026 study gave frontier agents the central open research question from two unpublished conference submissions, six days, and substantial compute. The agents completed the engineering but made no substantial progress on the research questions. A second model and scaffold reproduced the failures.

WTK's floor-raising claim is testable for goals with checkable acceptance obligations. It remains unsupported for goals whose acceptance criteria are primarily open-ended judgments. The evidence covers two papers from one lab and frontier agents available in mid-2026, so it identifies a counterexample class rather than a general ceiling.

Share an approach or counterexample
RANK 6
STATE · namedMOVES THROUGH · operational

Delegation Integrity

Can a team or swarm become unsafe without any individual agent violating its local policy?

Why it exists and what moves it +

Locally permitted actions can compose into a globally prohibited outcome. Delegation integrity therefore cannot be established by evaluating each agent in isolation.

Falsification signal: Every participant remains within its declared allowance while the combined workflow violates a global constraint, exceeds a budget, reconstructs restricted information, bypasses separation of duties, or produces an unauthorized outcome.

Evidence direction: Run adversarial composition tests covering collusion, capability splitting, confused deputies, cyclic delegation, constraint loss at handoffs, and unsafe fan-in.

Repair-loop integrity, C-16, is the same failure shape across time rather than across agents.

Share an approach or counterexample
RANK 7
STATE · boundedMOVES THROUGH · research

Adversary Non-Stationarity

Are the safety guarantees stated against an adversary that stands still?

Why it exists and what moves it +

A 2026 self-play red-teaming system reported more successful prompt-injection attacks than human red-teamers and transfer to held-out environments, defender models, and harnesses. A safety result must therefore name the adversary tier against which it was measured.

This tests portability assumptions, but it does not replace WTK's rule that every harness projection earns qualification independently. The work is self-reported by its originating lab, lacks independent replication, and targets prompt injection rather than the full honesty surface.

Share an approach or counterexample
RANK 8
STATE · boundedMOVES THROUGH · research

Repair-Loop Integrity

Can a sequence of individually compliant repairs arrive somewhere unsafe?

Why it exists and what moves it +

One 2026 paper identifies repair-time amplification in agent trajectories and constrains repair with provenance-preserving state transitions. Another shows how a solver and its rubric can co-adapt, making apparent progress come from an easier measurement rather than better work.

If WTK ever allows judged criteria to evolve, aggregate pass rates cannot drive that evolution. The deterministic floor remains fixed. The provenance result is from multimodal reasoning benchmarks, while the score-decoupling result reports 2.8% to 5.0% relative gains against its own baseline, so neither study alone proves long-term factory safety.

On 2026-08-14, a deterministic WTK comparison replayed five fixed persistent failure classes through instruction-only repair and typed escalation. Typed routing selected the declared repair layer in 5 of 5 cases and authorized no repeated instruction-only retries. This is routing evidence, not proof that the selected repairs improve live outcomes or remain safe across repeated improvement cycles.

Share an approach or counterexample
RANK 9
STATE · namedMOVES THROUGH · operational

Unknown Unknowns

How does the factory fail when the failure was not anticipated by its threat model, schemas, fixtures, or operators?

Why it exists and what moves it +

Unknown unknowns cannot be enumerated in advance. The architectural requirement is bounded impact, observable anomaly, safe degradation, and conversion of novel failures into durable tests.

Falsification signal: An unmodeled failure causes silent unsafe continuation, unbounded impact, corrupted evidence, or a false success verdict.

Evidence direction: Use mutation, chaos, red-team, dependency-failure, and novel interaction testing. Measure blast radius, anomaly detection, time-to-safe-state, evidence preservation, and time from discovery to reusable regression test.

Share an approach or counterexample
RANK 10
STATE · namedMOVES THROUGH · external

Economic Value

Does governed autonomous work create enough net value that organizations voluntarily continue using it?

Why it exists and what moves it +

Technical autonomy is insufficient. Governance overhead, evaluation, human review, latency, remediation, compute, failures, and incident response are part of the cost.

Falsification signal: The system can perform autonomous work but its total cost, delay, risk, or operator burden consistently exceeds the value of the work or a simpler alternative.

Evidence direction: Compare governed workflows against human and less governed baselines using net value, time-to-outcome, intervention rate, failure cost, and quality-adjusted throughput.

Share an approach or counterexample
THE LEDGER

Still real. Still load-bearing. Never quietly discarded.

The Ledger groups challenges by what prevents them from entering The Ten today.

5 CHALLENGES

In motion

A live line of attack exists and the challenge is already bounded.

STATE · boundedMOVES THROUGH · research

Intelligence-Substrate Independence

Can the same governance obligations be expressed, enforced, and evaluated if autonomous systems move from large language models to materially different intelligence architectures?

Why it exists and what moves it +

Candidate substrates include JEPA-style predictive world models, neuro-symbolic systems, memory-centric architectures, agent-native reasoning systems, deterministic decision systems, and architectures not yet invented.

The architectural claim is that WTK governs purpose, authority, actions, evidence, evaluation, and consequences rather than prompts, tokens, tool-calling conventions, or inference internals. The governing obligations may survive even when their current adapters, contracts, or enforcement mechanisms must change.

Falsification signal: A core governance obligation cannot be represented or enforced without LLM-specific constructs, or materially different substrates receive incompatible guarantees for reasons the architecture cannot expose and bound.

The claim is bounded but not yet instrumented because no held-out cross-substrate test exists.

Evidence direction: Apply one target-neutral goal, authority, action, evidence, and outcome contract to both an LLM-based agent and a materially different predictive or decision system. Compare enforceability, evidence completeness, failure semantics, and outcome accountability without assuming behavioral equivalence.

Share an approach or counterexample
STATE · boundedMOVES THROUGH · operational

Evidence Authenticity and Completeness

Can evidence be forged, replayed, fabricated, altered, selectively omitted, or detached from the action it claims to represent?

Why it exists and what moves it +

Authentic evidence can still be incomplete. A signed record is not sufficient if material actions occurred outside the observed boundary.

This challenge covers forgery, replay, substitution, truncation, and selective omission. It does not cover evidence that is authentic, complete, and false because its source was false. See C-13.

Falsification signal: A verifier accepts evidence for a run that did not occur, accepts evidence from another run or actor, fails to detect modification, or reports a complete record while a material action was omitted.

Evidence direction: Test cryptographic binding among identity, policy, artifact, runtime, tool receipt, time, and outcome. Attempt replay, truncation, reordering, substitution, equivocation, forged receipts, and log suppression.

Share an approach or counterexample
STATE · boundedMOVES THROUGH · operational

Long-Term Drift

Does the cost of re-earning evidence survive the pace of model releases?

Why it exists and what moves it +

Qualification is temporal. Evidence earned against one executor tuple or policy snapshot does not automatically survive change.

Falsification signal: Material drift occurs without invalidating evidence, triggering requalification, or reducing the system's claimed confidence.

Evidence direction: Operate long-duration workloads with controlled model, dependency, tool, policy, and personnel changes. Measure intervention rate, unsafe-action rate, stale-evidence duration, recovery, cost variance, and repeated-remediation loops.

On 2026-08-14, one deterministic local package revision preserved the prior source and downstream artifacts while rejecting stale construction and operator-approval evidence for the revised digest. This supports the source-edit-to-catalog invalidation seam only. A complete rebuild, retest, requalification, target compile, and catalog readmission remains open.

Share an approach or counterexample
STATE · boundedMOVES THROUGH · research

Cross-Harness Equivalence

Can identical goals receive equivalent governance guarantees across Claude, Codex, GPT-based systems, local runtimes, and future harnesses?

Why it exists and what moves it +

Equivalent governance does not require identical behavior. It requires that differences in enforcement be discovered, measured, and represented honestly.

Falsification signal: A projection claims a guarantee the target cannot enforce, or qualification evidence from one harness is treated as proof for another harness or execution form.

Evidence direction: Use the same contract and fixtures across multiple harnesses, record target capability declarations and enforcement receipts, and compare outcomes without averaging away weaker guarantees.

This asks whether defenders are equivalent across harnesses. C-15 asks whether attackers transfer across them.

Share an approach or counterexample
STATE · boundedMOVES THROUGH · research

Harness Erosion

Does the governance layer still add measurable capability as models absorb the scaffold?

Why it exists and what moves it +

A 2026 spatial-reasoning method raised an 8B model from 29.3% to 84.6% with tools and retained 73.8% after tool capabilities were internalized. The tool-free result remained about 11 points below the tool-assisted result.

Internalization can reduce the capability value of a harness, but it also removes the explicit tool-call receipt and weakens external observability. It does not eliminate all possible validity evidence. The study is limited to vision and perception tasks.

Share an approach or counterexample
2 CHALLENGES

Cheap to instrument

A small, well-understood measurement has not yet been performed.

STATE · namedMOVES THROUGH · operational

Discovery-Channel Coverage

Is the search for disconfirming evidence broad enough to find what exists?

Why it exists and what moves it +

A falsification program can only be contradicted by work that reaches it. The research radar uses a daily feed plus targeted sweeps, but its miss rate and source diversity have not been measured.

Share an approach or counterexample
STATE · namedMOVES THROUGH · operational

Intake Judgment

Who checks the findings that were discarded?

Why it exists and what moves it +

Research triage sits upstream of every finding in this program. Discarded items need an independent sample review so the same judgment does not both reject a source and certify that rejection as correct.

Share an approach or counterexample
4 CHALLENGES

External or deferred

Movement depends on external evidence or operational work outside the current research cycle.

STATE · namedMOVES THROUGH · external

Human Trust Calibration

When should an operator stop checking every action, and will human trust track actual reliability?

Why it exists and what moves it +

Both overtrust and undertrust are system failures. The target is calibrated reliance, not maximum operator confidence.

Falsification signal: Operators increasingly accept unsafe outputs, continue manual review after the system has earned bounded autonomy, misunderstand degraded guarantees, or cannot identify when evidence is insufficient.

Evidence direction: Measure intervention quality, override accuracy, automation bias, alert fatigue, time-to-detect, and trust calibration as system reliability changes.

Share an approach or counterexample
STATE · namedMOVES THROUGH · external

Organizational Acceptance

Can an enterprise approve, operate, audit, and remain accountable for a WTK target-state deployment?

Why it exists and what moves it +

Technical correctness does not guarantee institutional adoption. Ownership, liability, procurement, policy, incident response, workforce impact, and regulatory interpretation are separate constraints.

Falsification signal: The system cannot produce evidence, responsibility boundaries, operating procedures, or risk decisions that real security, legal, compliance, audit, and business owners can accept.

Evidence direction: Submit bounded deployments to real architecture, security, privacy, legal, compliance, procurement, and operational-readiness reviews. Preserve objections and rejected deployments as evidence.

Share an approach or counterexample
STATE · namedMOVES THROUGH · operational

Replacement and Protocol Independence

Can WTK's target-state architecture coexist with or be superseded by another governance framework without trapping implementations or invalidating portable evidence?

Why it exists and what moves it +

Replaceability is a desired property. WTK's target-state architecture should define portable protocols and evidence semantics rather than make one implementation the permanent authority.

Falsification signal: Governance artifacts require one product's runtime internals, an independent system cannot verify them, or replacing the research framework invalidates evidence for reasons unrelated to the governed behavior.

Evidence direction: Exchange contracts, policies, lineage, receipts, and qualification records with an independent implementation. Test migration, dual-verification, and disagreement handling.

The work is operational, but meaningful movement requires an independent implementation or reference protocol and is deferred from the current cycle.

Share an approach or counterexample
STATE · namedMOVES THROUGH · external

Civilization

If governed autonomous work succeeds technically and economically, can society adapt to its consequences?

Why it exists and what moves it +

Labor markets, institutions, governments, education, ownership, inequality, and public legitimacy are systems outside the technical architecture. The architecture can succeed while civilization struggles.

Falsification signal: A technically conformant factory produces systemic harms that its deployment and governance model cannot recognize, constrain, or assign accountability for.

Evidence direction: Treat social outcomes as external deployment evidence, not as proof of technical conformance. Include affected stakeholders and institutions in defining acceptable operating envelopes.

Share an approach or counterexample
CHANGE LOG

Movement remains visible.

Initial Ranked Baseline

Twenty-one stable challenge records were established. The Ten received its first rank order. No state promotion is claimed in this entry.