Distinguishing mirage-based from genuine visual reasoning

Establish reliable, operational criteria and detection methods that can distinguish mirage-based reasoning from genuine image-grounded reasoning in large multimodal models’ explanations and chain-of-thought traces during visual question answering and related multimodal tasks.

Background

The paper shows that frontier multimodal models often produce detailed visual descriptions and correct answers even when no image is provided, a behavior termed mirage reasoning. In many evaluations, the models’ reasoning traces in mirage-mode appear indistinguishable from those generated with real images.

Because the same models’ explanations look visually grounded regardless of image presence, the authors note that the boundary between mirage-based and true visual reasoning cannot be readily discerned from the generated justifications alone, creating a critical evaluation and safety gap.

References

The distinction between mirage-based and visual thinking is unclear

MIRAGE: The Illusion of Visual Understanding  (2603.21687 - Asadi et al., 23 Mar 2026) in Subsection: "The distinction between mirage-based and visual thinking is unclear" (Section: Mirages give the illusion of visual understanding)

The mechanistic reading that the cognition stream structurally severs the path along which deep reasoning drifts from the visual evidence is confirmed by no rigorous causal analysis and should be read as a structural design intention only; the current choices of perception anchor layer and cognition injection layers follow design considerations, and no layer-by-layer sweep or sensitivity analysis was performed, so this combination cannot be asserted to be optimal; every difference reported here is a point estimate from a single evaluation run without significance testing; and the out-of-domain evaluations use a scoring protocol different from the official leaderboards, string matching on some benchmarks and large-model scoring on others, so their absolute scores are not directly comparable with any leaderboard and the associated conclusions are restricted to the relative ranking of the four configurations under one scorer, with MMHal-Bench returning a null result on a small sample of non-COCO images that serves only as boundary evidence for domain-conditionality.

Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors  (2608.12746 - Bu, 13 Aug 2026) in Section 6, Conclusion and Future Work