MoVT: Adaptive Multimodal Reasoning
- MoVT is an adaptive visual reasoning paradigm that integrates text-based and visually-grounded modes within a single unified autoregressive model.
- It employs a two-stage training process using supervised cold-start with mode-specific prefixes followed by RL-based mode selection via AdaGRPO.
- Empirical results demonstrate enhanced performance across diverse benchmarks by effectively differentiating reasoning modes based on task context.
Searching arXiv for the cited MoVT and related visual-thought papers to ground the article in current literature. Search 1: Mixture-of-Visual-Thoughts paper. Mixture-of-Visual-Thoughts (MoVT) is an adaptive visual reasoning paradigm in which a multimodal model contains multiple reasoning modes within a single unified autoregressive model and selects the appropriate mode according to the image–question context (Li et al., 26 Sep 2025). In its canonical formulation, MoVT unifies text-based reasoning and visually-grounded reasoning rather than committing to a single chain-of-thought format. The central claim is that different visual tasks favor different inductive biases: text-based reasoning is strong for abstract reasoning, especially math-like tasks, while visually-grounded reasoning is better for object-centric tasks, can reduce hallucination, and is useful when precise localization matters, but is weak when concepts are abstract and cannot be grounded visually (Li et al., 26 Sep 2025). In adjacent work, the broader conceptual basis for MoVT is supplied by the notion of visual thoughts as intermediate, logic-driven cross-modal representations, and by modal-mixed reasoning systems that interleave language with latent visual embeddings (Cheng et al., 21 May 2025, Shao et al., 31 Jan 2026).
1. Definition and conceptual basis
MoVT is defined by the decomposition
where is the image, is the question, is the mode prefix, is the thinking process, and is the answer (Li et al., 26 Sep 2025). This formalization separates reasoning into two coupled problems: mode selection and mode-conditioned reasoning and answering . The formulation places mode choice inside the autoregressive generation process rather than treating it as a separate heuristic or an inference-time-only trick.
The need for MoVT is motivated by the complementarity of two reasoning modes. The paper distinguishes text-based reasoning, which reasons purely in language like standard CoT, from visually-grounded reasoning, which explicitly anchors reasoning to image regions, typically by emitting coordinates such as bounding boxes (Li et al., 26 Sep 2025). The argument is not that one mode should replace the other, but that no single reasoning mode dominates across all tasks. This directly motivates a model that can both learn multiple modes and differentiate among them.
A related conceptual strand comes from the “visual thoughts” perspective, which argues that multimodal chain-of-thought improves reasoning because it inserts intermediate representations that carry image information into the reasoning stream (Cheng et al., 21 May 2025). That paper states that “Visual thoughts are intermediate, logic-driven cross-modal representations that facilitate and accelerate multimodal reasoning within a unified perspective. By caching distilled visual information, they bridge raw pixels and linguistic rationales, enabling fast, context-aware access without reprocessing the image.” This suggests that MoVT is not merely a policy over output styles; it can also be understood as a policy over distinct forms of intermediate visual state.
2. MoVT in the visual-thought literature
The immediate background to MoVT is the attempt to unify different multimodal chain-of-thought formats. “Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought” argues that Textual-MCoT and Interleaved-MCoT are not fundamentally different reasoning paradigms; rather, both are expressions of visual thoughts (Cheng et al., 21 May 2025). It defines four visual-thought forms: Natural Language (N-LANG), Structured Language (S-LANG), Edited Image (E-IMG), and Generative Image (G-IMG). Their reported task affinities are differentiated: N-LANG is strongest for coarse perception, S-LANG is best for relation reasoning, E-IMG excels on detailed attribute reasoning, and G-IMG is best for iterative, multi-step reasoning (Cheng et al., 21 May 2025).
This broader literature is relevant because MoVT can be interpreted as a restricted but operational mixture over visual-thought expression types. In the 2025 MoVT paper, the two instantiated modes are text-based reasoning and grounded reasoning (Li et al., 26 Sep 2025). In the 2025 “Visual Thoughts” paper, the space of expressions is broader and includes both textual and image-form intermediates (Cheng et al., 21 May 2025). A plausible implication is that the two-mode MoVT formulation is one concrete point within a larger design space of visual-thought mixtures.
A second related direction is “Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings,” which presents modal-mixed CoT in which the model alternates between text tokens and compact latent visual embeddings (Shao et al., 31 Jan 2026). There, the full generation trajectory is written as
with the 0 functioning as internal visual sketches. That paper explicitly characterizes itself as a concrete realization of the broader Mixture-of-Visual-Thoughts idea. The comparison is informative: the 2025 MoVT system mixes explicit reasoning modes under prefix control, whereas the 2026 modal-mixed system interleaves language with latent visual states inside the VLM (Shao et al., 31 Jan 2026).
3. AdaVaR: supervised cold-start and RL-based mode induction
The framework used to train MoVT is AdaVaR: Adaptive Visual Reasoning (Li et al., 26 Sep 2025). AdaVaR has two stages. The first is a supervised cold-start stage that teaches the base model to produce multiple reasoning modes in a unified format. The second is an RL stage that induces context-dependent mode selection and further improves reasoning.
In the supervised stage, AdaVaR uses a uniform sequence format for all modes and assigns each mode a unique prefix token: <ground> for grounded reasoning and <text> for text-based reasoning (Li et al., 26 Sep 2025). The prompt instructs the model that grounded reasoning should output coordinates like object[x1, y1, x2, y2], while text-based reasoning should reason only in text, without localization. The model response is structured as mode prefix, > ..., and <answer> ... </answer>.
The SFT data mixes expert trajectories from both modes. The visually-grounded reasoning data is taken from high-quality SFT data from prior grounded reasoning work, especially VoCoT, and the final SFT setup reports a grounded dataset of 119K examples (Li et al., 26 Sep 2025). The text-based reasoning data is built by distilling a text-based visual reasoning model and applying rejection sampling; in the appendix, the paper states that it distills Orsta, sampling questions from R1-OneVision, LLaVA-CoT, NuminaMath1.5, and Virgo. After rejection sampling, it reports 115K reasoning examples and 95K direct-answer examples (Li et al., 26 Sep 2025).
A key claim is that simply mixing grounded and text-based data is not sufficient. The mode-specific prefix is described as crucial because it lets the model distinguish between reasoning modes, learn both under a unified interface, and later explore them evenly during RL (Li et al., 26 Sep 2025). This is one of the clearest ways MoVT differs from undifferentiated data mixing.
4. AdaGRPO and the mechanics of mode selection
The RL component of AdaVaR is AdaGRPO, a modification of GRPO designed for multi-mode reasoning (Li et al., 26 Sep 2025). The paper identifies two problems with vanilla GRPO in this setting: the model may under-explore modes and generate rollouts from only one mode, and rollout-level advantage alone does not explicitly compare modes.
AdaGRPO addresses this with three elements. First, it uses prefix-guided mode exploration: the 1 rollouts are split into 2 text-based rollouts with 3 and 4 grounded rollouts with 5, ensuring even exploration across modes (Li et al., 26 Sep 2025). Second, it uses a reward function following DeepSeek-R1-style rewards, with format rewards and accuracy rewards, and the accuracy reward is 1 or 0 based on rule-based evaluation of the produced answer. Third, it computes mode-relative advantages by modeling the rewards of the two mode-specific rollout sets as Gaussian distributions: 6 and then defining
7
or equivalently
8
as reported in the appendix (Li et al., 26 Sep 2025).
The token-level assignment is asymmetric by design. Mode prefix tokens receive the mode-relative advantage, while thinking tokens receive the rollout-level advantage 9 (Li et al., 26 Sep 2025). The intended effect is explicit: prefix tokens are optimized to encourage selecting the better mode, whereas reasoning tokens are optimized for the quality of the reasoning process itself.
The paper also states that RL includes a KL penalty with coefficient 0.04, and it uses curriculum learning with two data phases: a binary mixture containing only OmniCount and Geo170K, followed by a diverse mixture containing the remaining RL data, including math, OCR, object counting, science, grounding, and document tasks (Li et al., 26 Sep 2025). This curriculum is intended to help the model first learn coarse mode distinctions and then refine them on harder tasks.
5. Empirical performance and mode-selection behavior
The reported models are AdaVaR-3B built from Qwen2.5-VL-3B and AdaVaR-7B built from Qwen2.5-VL-7B (Li et al., 26 Sep 2025). Evaluation is conducted on eight benchmarks: MathVista, MathVision, MathVerse (vision-only split), WeMath, MMStar, V*, POPE, and SpatialScore. These benchmarks cover mathematical reasoning, general perceptual reasoning, visual search, spatial reasoning, and hallucination diagnosis.
For AdaVaR-3B, the paper reports: MathVista 69.8, MathVision 24.5, MathVerse 35.2, WeMath 33.8, MMStar 59.3, V* 77.0, POPE 88.2, SpatialScore 18.9, with average 50.84 (Li et al., 26 Sep 2025). For AdaVaR-7B, it reports: MathVista 74.4, MathVision 28.5, MathVerse 43.0, WeMath 44.8, MMStar 63.0, V* 83.4, POPE 89.0, SpatialScore 20.4, with average 55.82 (Li et al., 26 Sep 2025). The paper highlights that AdaVaR-3B matches or approaches much larger models, that AdaVaR-7B surpasses GPT-4o in average performance, and that AdaVaR is the only model family in its table that improves across all datasets relative to the base Qwen2.5-VL models (Li et al., 26 Sep 2025).
The behavioral evidence for adaptive mode selection is central. The paper reports that on math problems, AdaVaR tends to choose text-based reasoning; on V* and POPE, it tends to choose grounded reasoning; and on MMStar and other mixed tasks, it makes more nuanced category-dependent choices (Li et al., 26 Sep 2025). A specific example is given for MathVista: in SFT, grounded mode is chosen 31% of the time, whereas after RL, grounded mode choice drops to 2%. This is presented as a clear demonstration that RL is needed for mode selection.
The ablations support the design choices. Removing both Ada-Adv and PG-Exp, which is equivalent to standard GRPO, causes the model to reinforce the SFT model’s bias toward grounded mode and substantially reduces performance (Li et al., 26 Sep 2025). Removing Ada-Adv only weakens mode selection while retaining some benefit from prefix-guided exploration. Removing diverse mixed data or removing curriculum learning also degrades performance. Most notably, removing the mode prefix yields a “Mix-SFT-RL” variant that performs worse than even single-mode baselines, which the paper treats as evidence that explicit mode identity is essential (Li et al., 26 Sep 2025).
6. Interpretive significance, misconceptions, and limitations
One common misconception is that MoVT is merely a matter of combining heterogeneous training data. The MoVT paper explicitly argues otherwise: the model does not just “mix” data; it learns to differentiate the reasoning modes and adaptively choose among them (Li et al., 26 Sep 2025). Another misconception is that multimodal reasoning necessarily requires explicit image generation or external tools. Related work shows multiple possibilities: “Visual Thoughts” frames both textual and image-form intermediates as expressions of the same underlying mechanism (Cheng et al., 21 May 2025), whereas modal-mixed CoT uses latent visual embeddings generated and consumed inside the VLM rather than explicit rendered images (Shao et al., 31 Jan 2026).
The mechanistic interpretation offered by the visual-thought literature is that visual thoughts serve as intermediaries between the input image and deeper transformer layers (Cheng et al., 21 May 2025). Attention analysis in that work indicates that, when visual thoughts are present, attention shifts away from the raw image and toward the visual thought; information-flow analysis suggests that most image information first flows into the visual thought and then from the visual thought into the reasoning process. This suggests that MoVT’s adaptive mode selection can be viewed not only as selecting an output style but also as selecting how visual information should be compressed and transmitted during reasoning.
The current MoVT instantiation also has explicit limitations. The paper states that AdaVaR does not reach the upper bound of adaptive reasoning, that mode selection is not perfect, that grounded reasoning is limited on abstract tasks, that text-based reasoning can still hallucinate or overthink, that the current system focuses on English scenarios, and that grounded mode struggles in cases such as multi-image reasoning and abstract concepts that cannot be grounded (Li et al., 26 Sep 2025). It suggests future extensions including more reasoning modes beyond text and grounding, a direct-answer mode, search-based or non-linear reasoning modes, explicit reasoning about which mode to use, richer RL data, and multi-round SFT+RL iteration.
Taken together, these strands support a general interpretation of MoVT: general multimodal reasoning may require a mixture of reasoning styles rather than a single canonical chain-of-thought format (Li et al., 26 Sep 2025). In the two-mode AdaVaR formulation, the mixture is between text and grounding. In the visual-thought literature, the mixture extends across N-LANG, S-LANG, E-IMG, and G-IMG (Cheng et al., 21 May 2025). In modal-mixed latent reasoning, the mixture is between language tokens and latent visual sketches (Shao et al., 31 Jan 2026). The common principle is that reasoning quality depends on access to multiple representational formats and on learning when each format is appropriate.