WTKRESEARCH + ENGINEERING
← Back to Journal

Better Evidence Can Still Produce a Worse Decision

More grounded explanation does not guarantee that an agent will honor its hard constraints.

A decision can look better explained while becoming less correct where the rules are hardest. That is the uncomfortable result in a new multimodal benchmark: richer visual evidence improved several forms of reasoning, yet it made constrained resource allocation worse for every model tested.

The signal

In Seeing Is Not Deciding, the authors put nine frontier models through five executive decision tasks in 50 paired text-only and multimodal business scenarios. Multimodal input improved evidence-centric reasoning, with the largest reported gains in risk forecasting and board-facing justification. But adding visual business information degraded constrained resource allocation for all nine models.

That is a useful contradiction for agent builders. We often treat better grounding as a universal upgrade. This paper says the model may absorb more relevant evidence, describe it more persuasively, and still do worse at choosing an action that fits the stated limits.

The evidence

The paper's ablations point to signal crowding. Individual visual channels helped, but their combination disrupted constraint satisfaction during decoding. The authors call visual perception and constrained action separate bottlenecks.

Our take: an explanation score and an action-validity score should not be allowed to stand in for one another. An agent that cites the right materials or gives a sharper rationale has not thereby shown that it respected a budget, a permission boundary, a required sequence, or another non-negotiable condition.

The boundary

This is a controlled business-decision benchmark, not a test of WTK, tool execution, source provenance, or open-world research. It does not show that multimodal inputs are broadly harmful, that more sources are unsafe, or that one input format will transfer to every agent harness.

The result also does not supply a general effect size or a ready-made remedy. The reported mechanism is an ablation result under this benchmark's conditions. It gives us a strong hypothesis to test, not a maturity claim for WTK.

The builder impact

When an agent has hard constraints, measure two things separately: whether it used the evidence well and whether its final action conforms to every declared rule. Test the same task with each source channel alone and with their combination. Keep the success criteria independent enough that a fluent explanation cannot compensate for a forbidden or infeasible action.

This matters beyond images. A tool result, a retrieved document, a policy excerpt, and a long conversation history can all make the context richer. More context may improve an agent's account of the situation while making its decision surface harder to navigate.

The WTK test

WTK should treat grounding and constraint conformance as separate things to qualify. A useful paired fixture would hold the goal and hard rules fixed, vary the available evidence channels, and record both evidence use and rule-by-rule action validity. The result should remain target-specific: a compiled runtime earns its own proof instead of inheriting confidence from a more articulate response elsewhere.

WTK's grounding floor can still require evidence. The lesson is that evidence requirements need a companion check for the governed action that follows.

Still unknown

We do not yet know which combinations of sources create this failure mode, whether selective presentation can avoid it, or how it changes with tool-bearing agent tasks. The next experiment is modest: vary one evidence channel at a time, preserve the same constraints, and count every valid and invalid action before drawing a broader conclusion.

Have an approach, result, or counterexample?

You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.

Contribute to this research question
Working with an AI assistant?

Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.

Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.

RECORD DETAILSReference RN-009
Artifact
Research Notes
Status
Published
Evidence posture
Published with the evidence boundary stated in this record
Published
August 7, 2026
Author
WTK Research
Review
WTK human editorial review
Linked sources
1