WTKRESEARCH + ENGINEERING

Governed AI Agent Factory Architecture

Govern the environment, not the intelligence.

A governed AI agent factory separates package source truth, runtime execution, evidence, qualification, and promotion authority so the factory can automate more work without allowing capability to become authority.

What is autonomous agent factory governance?

Autonomous agent factory governance is the system of contracts, permissions, identity, evaluation gates, runtime controls, and evidence that governs how AI agents and agent teams are built, qualified, deployed, operated, and improved.

In the WTK architecture, the governing package remains outside the model. A cognition provider may propose and coordinate work, but it does not become the source of authority, define its own success criteria, or silently promote a replacement.

This separation makes each runtime projection independently inspectable: authority is bounded before execution, consequential actions are mediated, claims remain tied to evidence, and changes return through qualification and explicit promotion.

The factory supplies candidates. A separate authority governs replacement.

Improvement candidates originate in the factory, but promotion authority remains outside the candidate and the component that produced it. Today that authority is human-operated. Over time, bounded promotion decisions may be delegated to a separate policy-governed plane with explicit evidence requirements, risk limits, rollback conditions, and human escalation.

  1. 01Factory executionCandidate plus retained evidence
  2. 02Evaluation and qualificationScoped verdict plus limitations
  3. 03Separate promotion authorityHuman-authorized today
  4. 04Release and operationMonitor, suspend, revoke, or roll back
Explore governed self-improvement
WTK / TARGET-STATE STANDARD

Target-state pattern language

Research plates show proposed patterns, bounded observations, gaps, and falsification pressure.

Canonical packageThe authority-bearing source.
Derived projectionA harness-specific representation.
Enforcement gateA deterministic allow or block point.
Human decisionAn explicit operator judgment.
Evidence pathA receipt-backed observation.
Trust boundaryAuthority changes across this line.
01

Purpose and goals

Define the value sought, the boundaries that matter, and the conditions under which work must stop.

02

Governance contracts

Represent authority, permissions, delegation, escalation, and degraded behavior as inspectable structure.

03

Cognition and coordination

Allow models and agents to reason inside bounded roles without making the intelligence substrate the governing authority.

04

Evidence and evaluation

Bind claims to receipts, traces, deterministic checks, independent evaluation, and explicit uncertainty.

05

Qualification and promotion

Separate existence from readiness, semantic judgment from deterministic floors, and candidates from promoted assets.

06

Operation and learning

Observe outcomes, detect drift, degrade safely, and feed verified learning back into the architecture.

Six views of the same governed environment.

Each plate states the proposed pattern, bounded observations from an independent case study, what remains unproven, and which falsification challenge applies.

SOURCE-CONTROLLED ARCHITECTURE VIEW

Partially supported

Governed system context

Authority, package source truth, runtime execution, evidence, outcomes, and research remain distinct parts of the system.

WTK ARCHITECTURE / NATIVE READING VIEW
READING AXISAUTHORITY -> PACKAGE -> EXECUTION -> EVIDENCE -> LEARNING

Authority enters separately from cognition. The canonical package remains source truth while execution, evidence, outcomes, and research stay distinct.

  1. PrincipalGoal and delegated authority
  2. Factory controlGoverned construction
  3. Canonical packageContracts remain source truth
  4. Target projectionDerived for one execution form
  5. RuntimeOrchestration and cognition
  6. Evidence planeRun and qualification records
  7. ResearchBounded architecture learning
Authority and qualification

Governance supplies policy and promotion authority. Qualification returns a bounded verdict before projection.

Mediated consequence

Runtime exchanges bounded work with cognition, authorizes tools, and receives results and consequences.

Outcome return

Evidence reaches outcome stakeholders. Field evidence informs research, which may support, revise, narrow, or falsify guidance.

CONTROLLED RETURN PATH

Research returns bounded findings to governance; it does not become authority.

WTK IMPLEMENTATION OBSERVATION

WTK separates goal intake, canonical packages, qualification, target projection, runtime activity, and evidence records.

REMAINS UNPROVEN

Equivalent enforcement across intelligence substrates, runtime targets, and organizational boundaries has not been established.

FALSIFICATION PRESSURE

Intelligence-substrate independence · Cross-harness equivalence

SOURCE-CONTROLLED ARCHITECTURE VIEW

Architecture-defined

Trust boundary map

Authority, Factory, runtime, external consequence, and evidence domains require separate controls and identities.

WTK ARCHITECTURE / NATIVE READING VIEW
READING AXISINTENT -> FACTORY -> RUNTIME -> CONSEQUENCE -> VERIFICATION

Each column is a separate trust domain. Crossing a column boundary changes which identity, authority, control, or evidence must be checked.

  1. TRUST DOMAINAuthority domain
    • Principal
    • Governance and promotion authority
    • Identity and attestation authorities
  2. TRUST DOMAINFactory domain
    • Goal intake
    • Builder and package source
    • Protected evaluators and fixtures
    • Projection compiler
  3. TRUST DOMAINRuntime domain
    • Orchestration substrate
    • Cognition artifact
    • Governed context and memory
  4. TRUST DOMAINExternal consequence
    • Tool or service
    • External state
  5. TRUST DOMAINEvidence and consumer
    • Receipts and attestations
    • Independent consumer or verifier
    • Outcome stakeholders
BOUNDARY CROSSINGS

Each crossing changes the control or evidence obligation.

  1. IntentPrincipal -> goal intake
  2. Policy and evaluationGovernance -> builder -> protected evaluators
  3. Promotion and deploymentGovernance -> projection compiler -> substrate
  4. Identity and contextIdentity authorities and governed memory -> substrate
  5. AuthorizationAgent proposal -> substrate -> tool or service
  6. Receipt and verificationTool, substrate, and evaluator -> attestations -> independent verifier
WTK IMPLEMENTATION OBSERVATION

WTK represents contracts, protected evaluation, projection, identity context, receipts, and consumer-facing evidence as separate governed artifacts.

REMAINS UNPROVEN

Independent attestation, multi-operator boundaries, collusive-agent behavior, and external verifier interoperability remain open.

FALSIFICATION PRESSURE

Delegation integrity · Evidence authenticity

SOURCE-CONTROLLED ARCHITECTURE VIEW

Partially implemented

Consequential action authorization

Agents propose actions; an accountable substrate decides whether those actions may produce consequences.

WTK ARCHITECTURE / NATIVE READING VIEW
READING AXISBIND -> DELEGATE -> CHECK -> ALLOW, ESCALATE, OR DENY -> RECEIPT

The worker proposes a consequential action. The orchestration substrate owns the authorization decision and every outcome produces evidence.

PARTICIPANTS
  • Principal
  • Coordinator
  • Orchestration substrate
  • Worker workload
  • Tool or target
  • Evidence plane
  1. Bind the requestThe principal supplies the goal, identity, and authority ceiling. The coordinator receives a narrowed contract.
  2. Narrow delegationThe substrate verifies parent authority, then binds workload identity, task, context, and delegated authority.
  3. Authorize actionThe worker proposes an action. Identity, delegation chain, policy, audience, expiry, nonce, and budget are checked.
EXPLICIT OUTCOMES

Branches remain visible; no outcome is inferred from the primary path.

Allowed

A signed scoped envelope reaches the target. The target verifies consequence semantics; the result, receipt, and ordered attestations return.

Human decision required

A bounded approval request reaches the principal. The signed approve-or-deny decision is recorded in evidence.

Denied or degraded

The denial or degradation outcome is recorded and the worker receives a structured safe-state result.

WTK IMPLEMENTATION OBSERVATION

WTK defines narrowed role contracts, operator gates, structured denials, bounded handoffs, and attributable activity evidence.

REMAINS UNPROVEN

Target-side identity, nonce and replay enforcement, expiry, budget controls, and consequence semantics are not uniformly enforced across adapters.

FALSIFICATION PRESSURE

Evidence authenticity · Organizational acceptance

SOURCE-CONTROLLED ARCHITECTURE VIEW

Implemented

Governed artifact lifecycle

Creation, qualification, promotion, projection, deployment, operation, rollback, quarantine, and retirement are different states.

WTK ARCHITECTURE / NATIVE READING VIEW
READING AXISDEFINE -> BUILD -> QUALIFY -> PROMOTE -> PROJECT -> OPERATE

Lifecycle states are not interchangeable. Readiness, promotion, projection, deployment, operation, and retirement each require distinct transitions.

  1. DefineGoal draft -> goal accepted after the principal confirms intent and authority.
  2. BuildPackage candidate -> evaluation ready after governed construction and structural readiness.
  3. QualifyRequired evidence either passes into qualified or routes to remediation.
  4. PromoteAn authorized decision promotes a qualified package; withheld promotion leaves it qualified.
  5. Project and deployTarget compilation creates a projection; target validation and authorization permit deployment.
  6. OperateThe deployed artifact operates until remediation, quarantine, rollback, retirement, or invalidation changes its state.
Remediation return

Failed or incomplete evaluation and operational findings create a new package candidate version.

Anomaly controls

A critical anomaly quarantines the artifact; quarantine may route to remediation or retirement.

Rollback

A declared rollback restores a prior qualified version to operation.

Invalidation

Material source or policy change, evidence invalidation, or revoked promotion returns work to rebuild and requalification.

CONTROLLED RETURN PATH

Recovery paths return through governed construction and qualification; they never silently replace the incumbent.

WTK IMPLEMENTATION OBSERVATION

WTK separates candidate packages, evaluation readiness, qualification, promotion, deployment projection, invalidation, and remediation.

REMAINS UNPROVEN

Long-duration drift detection, revocation propagation, fleet rollback, and policy migration require sustained operational trials.

FALSIFICATION PRESSURE

Long-term drift · Human trust

SOURCE-CONTROLLED ARCHITECTURE VIEW

Partially supported

Evidence authenticity lifecycle

A trust claim is useful only when producer, authority, subject, scope, ordering, completeness, and threshold can be independently checked.

WTK ARCHITECTURE / NATIVE READING VIEW
READING AXISSUBJECT -> PRODUCER -> STATEMENT -> BINDING -> STORE -> CONSUMER

Evidence becomes usable only when its subject, producer, authority, ordering, storage, scope, completeness, and threshold can be checked independently.

  1. Evidence subjectArtifact, action, run, evaluation, or outcome
  2. Authorized producerProducer identity and authority are in scope
  3. Addressed statementContent is bound to its digest
  4. Identity bindingAuthority, time, and order are attached
  5. Evidence storeAppend-only or externally anchored
  6. Independent consumerChecks the record without trusting its producer
INDEPENDENT CONSUMER GATEEvery check must remain explicit.
  • Authentic producer?
  • Authorized claim?
  • Correct subject and scope?
  • Complete expected record?
  • Meets the declared threshold?
EXPLICIT OUTCOMES

Branches remain visible; no outcome is inferred from the primary path.

Applicable and sufficient

The consumer emits only the conclusion supported by the checked scope and threshold.

Rejected or inconclusive

Any failed check produces an incomplete, inconclusive, rejected, or invalidated result.

CONTROLLED RETURN PATH

Rejected evidence may be remediated or recollected by an authorized producer; the failed record remains visible.

WTK IMPLEMENTATION OBSERVATION

WTK emits structured run, audit, evaluator, tribunal, qualification, provenance, and digest-bound evidence records.

REMAINS UNPROVEN

Tamper-resistant storage, external anchoring, replay defense, selective-omission detection, and third-party verification remain incomplete.

FALSIFICATION PRESSURE

Evidence authenticity · Unknown unknowns

SOURCE-CONTROLLED ARCHITECTURE VIEW

Research active

Architecture learning loop

The architecture should change through bounded falsification evidence—not confidence, novelty, or silent self-promotion.

WTK ARCHITECTURE / NATIVE READING VIEW
READING AXISCLAIM -> CHALLENGE -> TEST -> IMPLEMENT -> EVIDENCE -> CONCLUDE

Architecture guidance changes through bounded falsification evidence. The implementation is an observation target, not the authority that decides what is true.

  1. Architecture claimState a falsifiable proposition
  2. Falsification challengeName how the claim could fail
  3. Bounded protocolDeclare conditions and measurements
  4. ImplementationWTK or another observation target
  5. All run evidencePreserve successes and failures
  6. Bounded conclusionStay inside the tested scope
EXPLICIT OUTCOMES

Branches remain visible; no outcome is inferred from the primary path.

Retain guidance

Current evidence supports keeping the bounded claim.

Revise or narrow

Evidence changes the conditions or permitted scope.

Invalidate or replace

The claim does not survive the challenge.

CONTROLLED RETURN PATH

Every conclusion creates the next stronger challenge and returns to falsification—not to a declaration of final proof.

WTK IMPLEMENTATION OBSERVATION

WTK records falsification challenges and research cycles while separating remediation, qualification, and promotion decisions.

REMAINS UNPROVEN

Governance self-qualification, long-running field evidence, unknown failure discovery, and safe promotion of governance changes remain open.

FALSIFICATION PRESSURE

Governance independence · Unknown unknowns

Behavior is governed through structure.

Prompts may guide behavior, but durable authority lives in contracts, registries, policies, gates, receipts, and evidence. A model can propose an action; the surrounding system decides whether that action is permitted, attributable, and promotable.

What the governance architecture must answer.

Who can authorize an autonomous agent?

Authority begins with an accountable principal, narrows through explicit delegation, and is checked again when an action could create an external consequence.

How is an AI agent qualified?

Qualification binds a verdict to the exact package, target projection, model, tools, policies, tests, and observed evidence. A result from one execution form does not automatically transfer to another.

Can an agent factory improve itself safely?

Evidence may inform a new candidate, but the candidate cannot redefine success or promote itself. The accepted version remains recoverable until external authority approves a separately qualified replacement.