A Release Claim Has to Remember the Misses
A deployment claim is only as honest as the complete evidence set it keeps.
Review update: September 5, 2026
The original account below describes the engineering boundary reported on August 6. It is not a claim that every provider, runtime, and interface path has since been checked. This update adds an example and narrows the scope without changing the original publication date.
One successful retry is not the whole record
Consider an illustrative three-attempt release check: one timeout, one invalid output, and one success. Reporting only the last result loses both the earlier failures and the effort needed to reach it. This example explains the requirement; it is not a measured WTK result.
Our later development review identified a separate accounting boundary: transport retries beneath the factory's provider wrapper could happen outside its visible retry policy. A complete top-level ledger does not, by itself, establish complete accounting of underlying calls. We should not read the original claim as proof that this lower layer was already covered.
What a reader can inspect today
On September 5, two mocked regression checks passed in the integration checkout: an HTTP 500 produced exactly one transport invocation, and a controlled deadline expiry produced a failed terminal request without a completed model-response event. Retained editorial receipts support these narrow checks, not a complete reconstruction of historical attempts, a released fix, or a reliability gain.
The original article supplies an engineering account, not a public attempt ledger or executable reproduction. The research method defines the stronger evidence standard. The replay log explains a different boundary: retaining a later review without replacing an earlier verdict.
A follow-up needs a sanitized trace connecting one requested operation to all transport attempts, failures, and authorized retries. Until that is available, this record supports the design reasoning and reported local change, not a quantified reliability or cost improvement.
Original account: August 6, 2026
A release claim can look persuasive even when it remembers only the cleanest run. That is not a release process. It is selective memory with better typography.
WTK already separates a package from a stronger target-specific readiness claim. This cycle tightened the evidence boundary underneath that distinction.
Current state
A package can be useful to discover, inspect, and compile without being ready to claim success in a particular execution form.
A stronger claim needs evidence tied to the package and conditions that actually ran. It also needs the unsuccessful, interrupted, and no-longer-current parts of the story to stay visible.
Changes
WTK now fixes the identity of a runtime-proof effort before the work begins. The identity covers the governed package, its projected execution form, the declared checks, and the conditions needed to interpret the result.
The system also establishes the complete set of planned attempts before execution. A run cannot quietly become the whole story just because it passed first. Interrupted work remains visible, and evidence from a materially changed condition becomes historical rather than supporting the new claim.
Required runtime checks, rollback checks, and operator approval still have to agree with the current claim. If they do not, the stronger release language remains blocked.
Failures observed
This change addresses several ordinary ways evidence can drift upward:
- keeping only a favorable retry after an earlier failure;
- treating a later package or runtime condition as if old evidence still applied;
- losing an interrupted attempt and calling the remaining record complete;
- allowing a changed control to inherit proof from an earlier configuration;
- treating a compiled package as enough evidence for a release claim.
Assumptions removed
We no longer assume that the latest successful run is the relevant evidence.
We no longer assume that unchanged-looking package source means every runtime condition is unchanged.
We no longer assume that an incomplete attempt set can support a complete deployment claim.
Evidence boundary
WTK can now represent a fixed proof identity, preserve a planned set of attempts, and fail closed when required current evidence is absent or mismatched.
That does not establish that people understand these distinctions, that every target can preserve them, or that a target-scoped runtime claim predicts a useful production outcome.
Research question advanced
This cycle advances the question of whether discovery can remain useful without allowing catalog presence to imply readiness. It supplies bounded mechanism evidence that a stronger catalog claim can require a complete, current evidence set instead of a favorable remembered run.
Next hypothesis
The boundary should survive material changes to the execution conditions. The next test will vary one declared runtime condition at a time and check that old evidence remains historical while a new claim stays blocked until its own planned evidence set is complete.
Have an approach, result, or counterexample?
You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.
Contribute to this research question →Working with an AI assistant?
Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.
Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.