Two Passing Roles Can Still Fail Together
A team is not reliable just because each role passed alone.
Two agent roles can both pass their own checks and still fail together when the same hard mission arrives. That matters because an independence assumption can quietly turn two weakly diverse lanes into a reliability story neither lane earned.
The signal
Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence tested that assumption in a preregistered evaluation of 18,000 two-agent handoff missions, scored by deterministic code. When two instances of the same model failed, they co-failed on 90.0% of the missions where either one failed. The reported association was strong: log odds ratio 6.66, with a 95% confidence interval of [6.38, 7.00], and phi 0.916.
That is the contradiction worth holding onto. Redundancy only helps if the routes to failure are meaningfully different. Counting two role-local passes as if they were independent can over-credit a team exactly when shared model behavior makes the same miss likely twice.
The evidence
The authors compared two-agent handoffs and report that substituting a different model reduced the association in all six tested contrasts. They also report a registered null for changing vendor when the model was already different. Their proposed certificate does not multiply component rates. It uses measured co-execution moments, then computes a finite-sample lower bound without assuming a dependence structure.
In their experiment, adding four measured moment functions, from 10 to 14, raised the certified floor from 0.2455 to 0.4116. That is not a magic number for every team. It is a useful demonstration that richer joint evidence can change the assurance claim, even when the individual components did not change.
The boundary
This is a preprint about the authors' two-agent mission set, not a WTK evaluation or a general reliability law for every team. It does not show that all shared-provider lanes will have the reported dependence, that a different provider creates sufficient diversity, or that this certificate is the right method for every WTK package. It also does not establish target-runtime parity, safety, or a WTK maturity increase.
The builder impact
Our take: a team contract should make it possible to ask a joint question, not just stack role scorecards. If a builder claims that fan-in, backup lanes, or a second reviewer improves reliability, the evaluation needs matched co-executions and retained joint outcomes. Otherwise, the claim is an extrapolation from component evidence.
The practical trap is familiar. Two lanes with different prompts and job titles can still share the same model, tools, context, and blind spots. Surface variety is not evidence of statistical diversity.
The WTK test
For WTK, we would treat this as a qualification question for a declared team and its compiled target. A bounded test would retain each lane's model and provider identity, package and target digests, matched task outcomes, and the joint failure pattern. It would compare the observed joint result with the independence estimate, then keep any reliability statement scoped to the evidence that actually supports it.
That gives WTK a useful containment rule: role-local qualification can support role-local claims. A team-level reliability claim needs team-level evidence.
Still unknown
We do not yet know which shared conditions most often create correlated failures in governed teams, or what minimum co-execution evidence is proportionate for a given deployment claim. The next experiment is not to import a certificate. It is to run a predeclared, package-bound comparison of shared and diversified lanes, retain every outcome, and see whether the joint evidence changes the conclusion.
Have an approach, result, or counterexample?
You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.
Contribute to this research question →Working with an AI assistant?
Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.
Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.