A Gate Failure Has to Name Its Own Layer
A blocked candidate is not automatically evidence that the candidate is the problem.
A candidate can fail a gate even when the candidate behavior is not what failed. That sounds fussy until an evaluation system turns a generator mistake or a scoring contradiction into a verdict about the thing under review.
Current state
WTK's candidate-admission path uses held-out evaluation as a limit. A candidate does not enter the next stage because a draft looks promising. It must pass a bounded review that creates test fixtures and checks declared evidence and fail-closed behavior.
Changes
One completed local review cycle compared two builds that began from the same declared work. Both were blocked, but not for the reason first suspected. In one path, fixture creation stopped before a usable evaluation could begin because a duplicate name was not resolved. In another completed evaluation, two equivalent evidence-limited refusals were assessed differently. One was marked as an accomplishment failure even though the evaluator's accompanying rationale described the refusal as correct and appropriately fail closed.
The comparison removed an early assumption that a configuration choice explained the blockage. The observed conditions belong to fixture generation and score interpretation, respectively. They are not evidence that the candidate itself could not meet its declared task.
Failures observed
The important failure here is attribution. A gate can be doing useful work while still reporting the wrong layer as the cause of a block. When that happens, a candidate-level failure label hides the next repair: distinguish an input-generation fault from an actual evaluation result, and keep a correct refusal from becoming a negative accomplishment score by accident.
Assumptions removed
We cannot treat every admission block as a statement about the candidate. A governed review needs to say whether the evidence failed to generate, the fixture failed to represent the obligation, the score disagreed with the declared rule, or the candidate failed the task. Those are different facts with different remedies.
Evidence
This is one completed development observation with a bounded comparison and a retained scoring discrepancy. It does not include a repeated controlled experiment, a repaired rerun, a public evidence package, or a claim about failure rates. It does not show that every gate can make this mistake, that either observed condition will recur, or that a repair will improve admission quality.
Next hypothesis
The next useful test is deliberately small. Hold the declared work and evaluation contract fixed, force a duplicate-name condition, and compare paired fail-closed refusals that should receive the same outcome. Retain every generation result, score, rationale, and gate decision. A result after repair would still be evidence about those fixed conditions, not a general reliability claim.
Have an approach, result, or counterexample?
You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.
Contribute to this research question →Working with an AI assistant?
Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.
Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.