Every run held unsafe promotion, but ambiguous scoring kept the memory comparison inconclusive.
Confidence
Provisional
Still unknown
Repeated-cell stability, broader goals and models, whether stronger conflicting history changes decisions, and a clean catalog-eligible positive control
Next experiment
Repeat the matched design with unambiguous action-boundary fields, a valid catalog-positive control, explicit native success predicates, and repeated cells
One model binding followed changed observations in every positive case, then failed the strict evidence-absence contract in every negative case.
Confidence
Provisional, based on 24 retained attempts under one frozen question and response contract
Still unknown
Whether the failure replicates across other questions, prompts, harnesses, model versions, and adversarial observations
Next experiment
Repeat the causal-pair protocol across additional questions, bindings, and harnesses while separating abstention, output conformance, and receipt-reference scoring.
A fixed comparison found that typed escalation routed every persistent failure correctly while instruction-only repair repeated the same move.
Confidence
Provisional, based on one deterministic five-case comparison against a declared routing oracle
Still unknown
Whether applying the selected repairs improves live agent outcomes, regression rate, cost, latency, or operator effort
Next experiment
Apply instruction-only and typed repairs to the same retained live failures and compare resolution, regression, honest holds, cost, latency, and operator action.
In four fixed cases, WTK's capability ladder preserved fit, cost consent, discovery, and construction boundaries that an installed-first policy missed.
Confidence
Provisional, based on one deterministic four-case comparison against declared decisions
Still unknown
Whether model-derived fit judgments, live discovery, provisioning, setup, and post-provisioning work preserve the same boundaries
Next experiment
Run an attended comparison with live registry candidates, unavailable connectors, explicit cost consent, and measured goal accomplishment after provisioning.