A Bounded Rating Scale Can Manufacture an Effect
A preregistered comparison can still be unable to identify the claim it reports.
Freezing an analysis protects it from convenient revision. It does not make the endpoint valid. When a rating scale clips values at its floor or ceiling, a clean-looking interaction can be created by unequal attenuation rather than the preference the experiment meant to measure.
This is a concrete sequel to A Precise Score Can Still Measure the Wrong Thing. That earlier note asked whether an evaluation's observations can distinguish the behavior named in its claim. This one isolates a specific collision: two different latent response patterns can produce the same bounded ratings, and unequal clipping can create the interaction the evaluator reports.
The signal
Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect derives this problem and tests it in a preregistered audit of a frozen tutoring judge. The useful WTK mechanism is an identifiability and saturation check before a bounded judge contrast is treated as evidence for a causal preference.
The evidence
The audit made 990 judge calls over 55 paired response contexts, three profile arms, two response poles, and three repetitions. Its registered primary overall-rating interaction was null at +0.085 on a five-point scale, with a 95 percent bootstrap interval from -0.167 to +0.353.
An exploratory productive-struggle interaction was +0.378. A construction with zero differential preference but floor censoring reproduced +0.321, or 85 percent, of that apparent effect. On 17 of 30 weak-stratum stimuli the low pole sat at the floor in every arm, causing the difference-in-differences calculation to collapse to the other pole's movement.
The boundary
This is one judge, one five-point rubric, six acid-mixture problems, and 55 stimuli from 23 source tutoring runs. It does not estimate a corrected latent effect or show that every bounded rubric is invalid. The study also reports pole-construction confounds, no profile-without-dialogue control, correlated competence labels, and no multiplicity correction for the full reported set.
The builder impact
WTK evaluator evidence should retain every per-item and per-pole rating, the scale definition, bound occupancy, saturation partitions, and exact evaluator binding. The earlier identifiability question therefore becomes an operational check for this endpoint: could floor or ceiling attenuation produce the same reported interaction under a zero-effect fixture? Preregistration, evidence sealing, and evaluator separation remain important, but they do not answer that question by themselves.
The WTK test
Compare WTK's current bounded-rubric analysis with one candidate that adds predeclared saturation partitions and an independently specified alternative outcome representation. Freeze positive and negative fixtures, evaluator, prompt, scale, repetitions, attempt order, and analysis before scoring.
Measure agreement with deterministic fixture truth, false effects under zero-effect censoring controls, sensitivity to real seeded effects, inconclusive rate, cost, and catalog decision. The candidate succeeds only if it reduces false effects without hiding real ones or weakening existing contract checks. Reporting a non-identifiable result as a finding is a failure. Correctly returning inconclusive is an acceptable finding, not a failure. This is a proposed WTK experiment, not authorization to change a gate.
Still unknown
We do not know whether any current WTK evaluator contrast is materially affected by censoring or whether the candidate diagnostic changes a decision.
Have an approach, result, or counterexample?
You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.
Contribute to this research question →Working with an AI assistant?
Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.
Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.