WTKRESEARCH + ENGINEERING

AI Agent Evaluation, Qualification, and Promotion Governance

Trust is a stack, not a badge.

A convincing output is not evidence of readiness or release authority. WTK separates machine-checkable requirements, semantic judgment, independent review, qualification, and promotion governance.

THE PROBLEM

One green status can hide several different questions.

APPROACH

Keep checks, judgment, independent review, qualification, and human authority separate.

SOLUTION

Every layer emits its own evidence and can narrow—but never silently upgrade—the next claim.

01

Deterministic checks

Did the required structure and behavior checks pass?

Role
Blocking floor
Produces
Passed, failed, or not assessable
02

Semantic evaluation

Does an independent evaluator find that the result meets qualities rules cannot fully express?

Role
Independent assessment
Produces
Pass, fail, inconclusive, or error
03

Independent review

Does risk justify several perspectives or judge bindings instead of one?

Role
Optional independent panel
Produces
Findings, quorum, split, or no quorum
04

Qualification

What may be claimed about this exact package and execution form?

Role
Scoped conclusion
Produces
Claim plus limitations and invalidation rules
05

Promotion authority

May this qualified candidate replace an incumbent, enter a release channel, or receive greater authority?

Role
Independent lifecycle decision
Produces
Promote, withhold, escalate, suspend, or revoke

Qualification states what was demonstrated. Promotion grants lifecycle authority.

Qualification binds a scoped conclusion to an exact subject and execution form. Promotion decides whether that subject may replace an incumbent, enter a release channel, or receive greater authority. The candidate and the component that produced it do not control that decision.

Human accountability does not require a human click for every mature, low-risk action. It requires humans to define the authority envelope, approve consequential policy, receive material escalations, and retain suspension and revocation power.

A trust claim should carry its basis.

A receipt binds the claim to the exact subject, execution conditions, evidence, review, limitations, and integrity record. It is a forensic report—not a decorative badge.

SUBJECT

Exact package + execution form

PROJECTION

Model + harness + tools

SCOPE

Declared fixtures + conditions

CHECKS

Deterministic evidence

REVIEW

Semantic findings + dissent

CLAIM

Qualification + limits

INTEGRITY

Lineage + evidence digest

Multiple reviewers are useful only when their independence is real.

Use additional perspectives when one evaluator's bias or instability creates material risk. Technical references reserve the term tribunal for a formally composed review panel.

PERSPECTIVE REVIEW

Several lenses, different questions

Security, reliability, evidence sufficiency, goal alignment, and operational risk can produce distinct findings for remediation.

Best for broader failure discovery.
JUDGE QUORUM

Several judges, the same question

Independent bindings vote on the same bound evidence and an M-of-N policy reports pass, fail, split, inconclusive, or error.

Intended to expose single-judge variance in higher-risk decisions; the measurable benefit remains under test.