WTKRESEARCH + ENGINEERING
← Back to Journal

A Better Claim Needs Its Own Definition

An improvement score is only meaningful when its intended outcome is declared before the comparison begins.

Current state

An improvement comparison needs to answer a basic question before it can score anything: better at what? A platform-wide score can be useful, but it can also hide the specific outcome an operator is trying to improve for one agent or team.

Changes

One completed engineering cycle added a source-controlled way to declare the claims that define improvement for a particular comparison. The selected definition is frozen with the goal, baseline, candidate, evaluation views, and method before evidence is collected.

The resulting claim observations remain separate rather than being flattened into one universal score. Existing goal, safety, authority, evidence, output, reliability, and team-obligation gates still take precedence over a favorable claim result.

Failures observed

The cycle addressed an attribution gap, not a demonstrated outcome failure. A comparison could otherwise report a positive number without making clear which declared outcome that number represented, or whether an evaluator viewpoint had been averaged away.

Assumptions removed

We cannot assume that one inherited metric expresses every valid improvement goal. We also cannot let a candidate redefine success after its evidence has arrived, or let a favorable specialized claim bypass a governing safety or authority gate.

Evidence

The completed cycle introduced bounded improvement definitions, froze the chosen definition into the comparison record, and checked that malformed definitions or definitions that attempt to widen lifecycle authority are rejected. It also preserved the earlier comparison form through an explicit compatible default.

Limitations

This is a deterministic engineering observation, not evidence that a chosen definition measures the right outcome, that evaluators interpret a claim consistently, or that a candidate improves real work. It does not establish cross-target behavior, promotion readiness, or general reliability.

Next hypothesis

Use one declared improvement definition in a separately authorized comparison. Hold the package pair, evaluation method, attempt policy, and governing gates fixed. Retain every claim result and disagreement, then ask whether the definition remains useful without allowing it to relax the other boundaries.

Have an approach, result, or counterexample?

You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.

Contribute to this research question
Working with an AI assistant?

Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.

Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.

RECORD DETAILSReference FL-021
Artifact
Factory Logs
Status
Published
Evidence posture
Published with the evidence boundary stated in this record
Published
August 25, 2026
Author
WTK Research
Review
WTK human editorial review
Linked sources
None declared