A Better Claim Needs Its Own Definition
An improvement score is only meaningful when its intended outcome is declared before the comparison begins.
Current state
An improvement comparison needs to answer a basic question before it can score anything: better at what? A platform-wide score can be useful, but it can also hide the specific outcome an operator is trying to improve for one agent or team.
Changes
One completed engineering cycle added a source-controlled way to declare the claims that define improvement for a particular comparison. The selected definition is frozen with the goal, baseline, candidate, evaluation views, and method before evidence is collected.
The resulting claim observations remain separate rather than being flattened into one universal score. Existing goal, safety, authority, evidence, output, reliability, and team-obligation gates still take precedence over a favorable claim result.
Failures observed
The cycle addressed an attribution gap, not a demonstrated outcome failure. A comparison could otherwise report a positive number without making clear which declared outcome that number represented, or whether an evaluator viewpoint had been averaged away.
Assumptions removed
We cannot assume that one inherited metric expresses every valid improvement goal. We also cannot let a candidate redefine success after its evidence has arrived, or let a favorable specialized claim bypass a governing safety or authority gate.
Evidence
The completed cycle introduced bounded improvement definitions, froze the chosen definition into the comparison record, and checked that malformed definitions or definitions that attempt to widen lifecycle authority are rejected. It also preserved the earlier comparison form through an explicit compatible default.
Limitations
This is a deterministic engineering observation, not evidence that a chosen definition measures the right outcome, that evaluators interpret a claim consistently, or that a candidate improves real work. It does not establish cross-target behavior, promotion readiness, or general reliability.
Next hypothesis
Use one declared improvement definition in a separately authorized comparison. Hold the package pair, evaluation method, attempt policy, and governing gates fixed. Retain every claim result and disagreement, then ask whether the definition remains useful without allowing it to relax the other boundaries.
Have an approach, result, or counterexample?
You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.
Contribute to this research question →Working with an AI assistant?
Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.
Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.