WTKRESEARCH + ENGINEERING
← Back to Journal

Proposed Evidence Checks for Agent Teams

A proposal to account for every required team role before assessing the executed workflow.

Editorial clarification: September 8, 2026

The display title and summary now describe the proposed evidence checks directly. In the historical text, a lane means a required role's part of the workflow. Proof refers to a proposed execution-evidence record, not a universal guarantee of team reliability. The September 5 clarification and original proposal remain unchanged below.

Classification clarification: September 5, 2026

This record remains a design proposal. Its original Factory Log identifier and August 30 publication date are preserved for citation continuity; they do not establish a completed engineering cycle. Under our current editorial criteria, a new article containing only this proposal would be a Field Note.

No completed team-runtime study is added by this update. The original proposal follows below.

A small testable example

Consider an illustrative team with an extractor, a checker, and a summarizer. A plausible summary alone cannot tell us whether the checker ran or whether the extractor respected its input boundary.

Case to test Evidence the proposed assessor would need Intended interpretation
All required roles complete Matching role outputs, handoffs, bindings, and budgets Eligible for further evaluation, not automatic qualification
Checker is missing A visible missing required role Insufficient evidence for complete-team execution
A handoff belongs to an older bundle Identity mismatch retained with the attempt Old work cannot satisfy the current team claim
Budget evidence is absent An explicit gap, not an inferred success Required proof remains incomplete

Scroll horizontally to see every column.

These are proposed cases and expected interpretations, not observed results. Passing such checks would establish only the tested accounting boundary, not that the final answer is useful or correct.

The existing experiment recommendation retains its identity. Its execution specification must compare current WTK assessment with the proposed complete-team evidence check on the same cases, including a complete control. Broken examples alone would test the candidate against its own expectations, not show whether it improves WTK. No experiment or completed-runtime result is authorized by this clarification.

The release-evidence log explains complete-attempt accounting. The research method defines the required separation between a proposed check and an experimental result.

Original proposal: August 30, 2026

Current state

A team contract can describe the lanes, handoffs, and authority a target should preserve. A generated target bundle can preserve that description. Neither fact shows that the target ran the complete team or that the required work reached its declared completion boundary.

In practical terms, an agent team is not proven because its files compile or one specialist succeeds. Every required role, handoff, output, permission binding, and budget must be accounted for in the executed workflow. Consider a simple coordinator to researcher to reviewer team: the final answer is not a team success if the researcher never ran, the reviewer rejected the evidence, or the coordinator exceeded the authority or budget declared for the job.

That is the distinction between describing a team and qualifying its exact execution form in WTK's agent-building lifecycle.

Proposed proof boundary

A target-runtime proof should bind the complete team bundle to a reviewed set of local bindings, then retain an outcome for each required lane, each lane's output-contract check, and each declared budget. Those records must identify the exact package, contract, bundle, bindings, workflow, and assessment that support the claim.

The assessor should treat the team as one subject. A completed install or a successful role in isolation is not enough. The resulting claim should require matching package, contract, bundle, binding-review, lane, and budget evidence for the exact executed workflow.

The public WTK evidence ledger shows why those identities matter: an implementation state and an evidence-supported claim are different things.

Failure modes to test

The test must make incomplete states visible rather than presenting them as a team success. A missing or failed required lane, an output that misses its contract, a stale binding review, or an unmatched budget receipt should leave the result at compile-only. Those conditions may still be useful diagnostic evidence, but they should not support the stronger runtime claim.

Assumptions removed

We cannot infer a team-level proof from a valid bundle, a successful install, or separate role-level checks. The claim needs an account of the required workflow as it ran: which lanes completed, what they handed forward, whether their contracts held, and whether the executed bindings stayed within the declared limits.

Evidence required for a completed-runtime claim

This published design proposal does not include an evidence digest, source record, or sanitized receipt set for a completed target-runtime cycle. A future completed-runtime claim must attach and check those identities. Until then, the record must not claim that WTK has demonstrated the proposed boundary.

Limitations

This does not establish that the proposed assessor works, that teams are generally reliable, that a target preserves every topology, that the proof transfers to another runtime, or that either role would succeed on other goals. It does not qualify a package for publication or promotion.

Next hypothesis

Run matched target cases that deliberately break one required lane, reviewed binding, output contract, or budget receipt at a time. Retain every result with sanitized, independently checkable identities and verify that the assessor rejects each incomplete variant while preserving a valid control. That would test the rejection boundary, not general team quality. The test should follow WTK's research and falsification method and remain a proposal until a human approves the frozen protocol.

Have an approach, result, or counterexample?

You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.

Contribute to this research question
Working with an AI assistant?

Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.

Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.

RECORD DETAILSReference FL-023
Artifact
Factory Logs
Status
Published
Evidence posture
Historical proposal; no completed-runtime result claimed
Published
August 30, 2026
Author
WTK Research
Review
WTK human editorial review
Linked sources
None declared