Distinguishing mirage-based from genuine visual reasoning
Establish reliable, operational criteria and detection methods that can distinguish mirage-based reasoning from genuine image-grounded reasoning in large multimodal models’ explanations and chain-of-thought traces during visual question answering and related multimodal tasks.
References
The distinction between mirage-based and visual thinking is unclear
The mechanistic reading that the cognition stream structurally severs the path along which deep reasoning drifts from the visual evidence is confirmed by no rigorous causal analysis and should be read as a structural design intention only; the current choices of perception anchor layer and cognition injection layers follow design considerations, and no layer-by-layer sweep or sensitivity analysis was performed, so this combination cannot be asserted to be optimal; every difference reported here is a point estimate from a single evaluation run without significance testing; and the out-of-domain evaluations use a scoring protocol different from the official leaderboards, string matching on some benchmarks and large-model scoring on others, so their absolute scores are not directly comparable with any leaderboard and the associated conclusions are restricted to the relative ranking of the four configurations under one scorer, with MMHal-Bench returning a null result on a small sample of non-COCO images that serves only as boundary evidence for domain-conditionality.