WTKRESEARCH + ENGINEERING

FALSIFICATION PROGRAM

Try to prove it wrong.

Start with the largest unresolved challenges to the architecture. Open a record only when you want its evidence, limits, and related questions.

21 shown

THE TEN

Largest unresolved challenges.

Ranked by how strongly each question could change the architecture. Open a record for the reasoning and evidence boundary.

RANK 1
STATE · namedMOVES THROUGH · operational

Falsifiability in practice

What finding would force a change to the architecture, rather than a narrower claim?

Why it exists and what moves it +

Recent review cycles produced disconfirming evidence and narrower boundary statements, but no precommitted condition determined when the architecture itself must change. Narrowing a claim is legitimate work, but a program that can always shrink its promise cannot expose architectural failure.

This challenge remains first until a trip-wire is written in advance that names a class of result capable of changing the architecture itself.

RANK 2
STATE · namedMOVES THROUGH · operational

Yardstick validity

Can the measurements that grade these challenges be trusted?

Why it exists and what moves it +

Every state on this page is downstream of evaluation. Held-out sets do not yet have to disclose contamination controls, and the stability of a verifier's repeated judgments has not been audited.

This is separate from authority over the qualifier. A legitimate authority can still hold a crooked ruler.

RANK 3
STATE · boundedMOVES THROUGH · research

Input truth

Can a system with structurally perfect grounding still produce confidently wrong work?

Why it exists and what moves it +

A 2026 study introduced one credible but misleading document into deep-research sources and raised mean report-level false-conclusion adoption from 0% to 54.7% across the studied configurations. Source authority and presentation mattered, while search rank and additional documents had limited influence.

WTK is designed to provide structured traceability to the sources recorded for a bounded execution; it does not establish source truth. Cross-model verification alone did not close the studied gap, but this result does not establish that every independent evaluation design will fail. The study covers selected agents, backbones, and constructed misleading documents.

Narrowing relationshipPartially narrows C-05.

RANK 4
STATE · namedMOVES THROUGH · operational

Governance independence

Who qualifies the qualifier, evaluates the tribunal, and safely promotes governance changes?

Why it exists and what moves it +

A governance system cannot recurse forever. WTK must identify a terminal trust root, separation of duties, and the evidence required to change either one.

Authority over the qualifier and validity of its measurements are separate failures. Conflating them hides both.

RANK 5
STATE · boundedMOVES THROUGH · research

Goal-class boundedness

Is there a class of goals where governance adds cost and no floor?

Why it exists and what moves it +

A 2026 study gave frontier agents the central open research question from two unpublished conference submissions, six days, and substantial compute. The agents completed the engineering but made no substantial progress on the research questions. A second model and scaffold reproduced the failures.

WTK's floor-raising claim is testable for goals with checkable acceptance obligations. It remains unsupported for goals whose acceptance criteria are primarily open-ended judgments. The evidence covers two papers from one lab and frontier agents available in mid-2026, so it identifies a counterexample class rather than a general ceiling.

RANK 6
STATE · namedMOVES THROUGH · operational

Delegation integrity

Can a swarm become unsafe while every individual agent appears compliant?

Why it exists and what moves it +

Locally permitted actions can compose into a globally prohibited outcome. Qualification must therefore test the team as a system, not only its members in isolation.

Repair-loop integrity is the same failure shape across time rather than across agents.

RANK 7
STATE · boundedMOVES THROUGH · research

Adversary non-stationarity

Are the safety guarantees stated against an adversary that stands still?

Why it exists and what moves it +

A 2026 self-play red-teaming system reported more successful prompt-injection attacks than human red-teamers and transfer to held-out environments, defender models, and harnesses. A safety result must therefore name the adversary tier against which it was measured.

This tests portability assumptions, but it does not replace WTK's rule that every harness projection earns qualification independently. The work is self-reported by its originating lab, lacks independent replication, and targets prompt injection rather than the full honesty surface.

RANK 8
STATE · boundedMOVES THROUGH · research

Repair-loop integrity

Can a sequence of individually compliant repairs arrive somewhere unsafe?

Why it exists and what moves it +

One 2026 paper identifies repair-time amplification in agent trajectories and constrains repair with provenance-preserving state transitions. Another shows how a solver and its rubric can co-adapt, making apparent progress come from an easier measurement rather than better work.

If WTK ever allows judged criteria to evolve, aggregate pass rates cannot drive that evolution. The deterministic floor remains fixed. The provenance result is from multimodal reasoning benchmarks, while the score-decoupling result reports 2.8% to 5.0% relative gains against its own baseline, so neither study alone proves long-term factory safety.

RANK 9
STATE · namedMOVES THROUGH · operational

Unknown unknowns

Can the architecture fail safely when it encounters a failure nobody anticipated?

Why it exists and what moves it +

Unknown failures cannot be enumerated in advance. The test is whether an unmodeled event remains bounded, observable, recoverable, and convertible into a durable regression test.

RANK 10
STATE · namedMOVES THROUGH · external

Economic value

Does governed autonomous work create enough value that organizations voluntarily retain it?

Why it exists and what moves it +

Technical autonomy is insufficient if governance cost, latency, review burden, failure cost, or incident response consumes the value produced. This remains in The Ten because it is existential and can only move through real organizational use.

THE LEDGER

Still real. Still load-bearing. Never quietly discarded.

The Ledger groups challenges by what prevents them from entering The Ten today.

5 CHALLENGES

In motion

A live line of attack exists and the challenge is already bounded.

STATE · boundedMOVES THROUGH · research

Intelligence-substrate independence

Can the same governance obligations be expressed, enforced, and evaluated across LLM agents and materially different intelligence architectures?

Why it exists and what moves it +

WTK claims to govern purpose, authority, actions, evidence, evaluation, and consequences rather than prompts or tokens. JEPA-style predictive world models, neuro-symbolic systems, and deterministic decision systems provide concrete alternatives against which that boundary can be tested.

The claim is bounded but not yet instrumented because no held-out cross-substrate test exists.

STATE · boundedMOVES THROUGH · operational

Evidence authenticity

Can evidence, receipts, or provenance be forged, replayed, or selectively omitted?

Why it exists and what moves it +

This challenge covers forgery, replay, substitution, truncation, and selective omission. It does not cover evidence that is authentic, complete, and false because its source was false.

STATE · boundedMOVES THROUGH · operational

Long-term drift

Does the cost of re-earning evidence survive the pace of model releases?

Why it exists and what moves it +

WTK evidence applies to a pinned combination of contract, model, harness, and policy. Material change triggers requalification, making drift a recurring cost rather than only a slow-decay risk.

STATE · boundedMOVES THROUGH · research

Cross-harness equivalence

Can different model and agent harnesses provide equivalent governance guarantees?

Why it exists and what moves it +

Equivalent governance does not require identical behavior. It requires each projection to expose what it can enforce and to earn qualification independently without borrowing evidence from another execution form.

This asks whether defenders are equivalent across harnesses. Adversary non-stationarity asks whether attackers transfer across them.

STATE · boundedMOVES THROUGH · research

Harness erosion

Does the governance layer still add measurable capability as models absorb the scaffold?

Why it exists and what moves it +

A 2026 spatial-reasoning method raised an 8B model from 29.3% to 84.6% with tools and retained 73.8% after tool capabilities were internalized. The tool-free result remained about 11 points below the tool-assisted result.

Internalization can reduce the capability value of a harness, but it also removes the explicit tool-call receipt and weakens external observability. It does not eliminate all possible validity evidence. The study is limited to vision and perception tasks.

2 CHALLENGES

Cheap to instrument

A small, well-understood measurement has not yet been performed.

STATE · namedMOVES THROUGH · operational

Discovery-channel coverage

Is the search for disconfirming evidence broad enough to find what exists?

Why it exists and what moves it +

A falsification program can only be contradicted by work that reaches it. The research radar uses a daily feed plus targeted sweeps, but its miss rate and source diversity have not been measured.

STATE · namedMOVES THROUGH · operational

Intake judgment

Who checks the findings that were discarded?

Why it exists and what moves it +

Research triage sits upstream of every finding on this page. Discarded items need an independent sample review so the same judgment does not both reject a source and certify that rejection as correct.

4 CHALLENGES

External or deferred

Movement depends on external evidence or operational work outside the current research cycle.

STATE · namedMOVES THROUGH · external

Human trust

Where is the equilibrium between automation overtrust and destructive undertrust?

Why it exists and what moves it +

Both overtrust and undertrust are system failures. This challenge moves through observed operator behavior as reliability and evidence quality change.

STATE · namedMOVES THROUGH · external

Organizational acceptance

Can an enterprise approve a governed autonomous factory operationally as well as technically?

Why it exists and what moves it +

Technical controls do not settle ownership, liability, procurement, incident response, workforce impact, or institutional accountability. Real review is required.

STATE · namedMOVES THROUGH · operational

Replacement

Can a better governance framework supersede the current target-state architecture without trapping implementations?

Why it exists and what moves it +

Replaceability is a desired property, but meaningful movement requires an independent implementation or reference protocol against which contracts, evidence, and migration can be tested. That work is operational but deferred from the current cycle.

STATE · namedMOVES THROUGH · external

Civilization

If governed autonomy succeeds, can institutions, labor markets, and society adapt?

Why it exists and what moves it +

The architecture can succeed technically while society struggles with its effects. Movement requires evidence from affected institutions and communities rather than stronger software claims.

CHANGE LOG

Movement remains visible.

Initial ranked baseline

Twenty-one stable records established. The Ten received its first rank order. No state promotion is claimed in this entry.