Does RL post-training in MLLMs truly leverage visual information
Determine whether reinforcement learning–based post-training for Multimodal Large Language Models (such as Qwen2.5-VL) truly enables the models to learn from and utilize visual information in the training inputs, rather than primarily strengthening internal text-based reasoning patterns.
References
Although many studies have reported improved performance, it remains unclear whether RL training truly enables models to learn from visual information.
— Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models
(2604.03179 - Zhang et al., 3 Apr 2026) in Abstract (page 1)
Our training necessarily encourages the model to discover effective solutions through exploration, raising the question of whether this process also improves its evidence. Maybe the evidence capability does not improve.
— MR-IQA-2: Faithful Image Quality Reflection via Fine-Grained Credit Assignment
(2608.18579 - Li et al., 19 Aug 2026) in Section 5, Discussion, paragraph “Good Evidence, Bad Solution?”