Keep checks, judgment, independent review, qualification, and human authority separate.
AI Agent Evaluation, Qualification, and Promotion Governance
Trust is a stack, not a badge.A convincing output is not evidence of readiness or release authority. WTK separates machine-checkable requirements, semantic judgment, independent review, qualification, and promotion governance.
One green status can hide several different questions.
Every layer emits its own evidence and can narrow—but never silently upgrade—the next claim.
Deterministic checks
Did the required structure and behavior checks pass?
- Role
- Blocking floor
- Produces
- Passed, failed, or not assessable
Semantic evaluation
Does an independent evaluator find that the result meets qualities rules cannot fully express?
- Role
- Independent assessment
- Produces
- Pass, fail, inconclusive, or error
Independent review
Does risk justify several perspectives or judge bindings instead of one?
- Role
- Optional independent panel
- Produces
- Findings, quorum, split, or no quorum
Qualification
What may be claimed about this exact package and execution form?
- Role
- Scoped conclusion
- Produces
- Claim plus limitations and invalidation rules
Promotion authority
May this qualified candidate replace an incumbent, enter a release channel, or receive greater authority?
- Role
- Independent lifecycle decision
- Produces
- Promote, withhold, escalate, suspend, or revoke
Qualification states what was demonstrated. Promotion grants lifecycle authority.
Qualification binds a scoped conclusion to an exact subject and execution form. Promotion decides whether that subject may replace an incumbent, enter a release channel, or receive greater authority. The candidate and the component that produced it do not control that decision.
Human accountability does not require a human click for every mature, low-risk action. It requires humans to define the authority envelope, approve consequential policy, receive material escalations, and retain suspension and revocation power.
A trust claim should carry its basis.
A receipt binds the claim to the exact subject, execution conditions, evidence, review, limitations, and integrity record. It is a forensic report—not a decorative badge.
Exact package + execution form
Model + harness + tools
Declared fixtures + conditions
Deterministic evidence
Semantic findings + dissent
Qualification + limits
Lineage + evidence digest
Multiple reviewers are useful only when their independence is real.
Use additional perspectives when one evaluator's bias or instability creates material risk. Technical references reserve the term tribunal for a formally composed review panel.
Several lenses, different questions
Security, reliability, evidence sufficiency, goal alignment, and operational risk can produce distinct findings for remediation.
Best for broader failure discovery.Several judges, the same question
Independent bindings vote on the same bound evidence and an M-of-N policy reports pass, fail, split, inconclusive, or error.
Intended to expose single-judge variance in higher-risk decisions; the measurable benefit remains under test.