---
title: 'MoVT: Adaptive Multimodal Reasoning'
url: https://www.emergentmind.com/topics/mixture-of-visual-thoughts-movt
type: topic
---

# MoVT: Adaptive Multimodal Reasoning

Searching arXiv for the cited MoVT and related visual-thought papers to ground the article in current literature.
Search 1: Mixture-of-Visual-Thoughts paper.
Mixture-of-Visual-Thoughts (MoVT) is an adaptive visual reasoning paradigm in which a multimodal model contains multiple reasoning modes within a single unified autoregressive model and selects the appropriate mode according to the image–question context [2509.22746]. In its canonical formulation, MoVT unifies text-based reasoning and visually-grounded reasoning rather than committing to a single chain-of-thought format. The central claim is that different visual tasks favor different inductive biases: text-based reasoning is strong for abstract reasoning, especially math-like tasks, while visually-grounded reasoning is better for object-centric tasks, can reduce hallucination, and is useful when precise localization matters, but is weak when concepts are abstract and cannot be grounded visually [2509.22746]. In adjacent work, the broader conceptual basis for MoVT is supplied by the notion of *visual thoughts* as intermediate, logic-driven cross-modal representations, and by modal-mixed reasoning systems that interleave language with latent visual embeddings [2505.15510; 2602.00574].

## 1. Definition and conceptual basis

MoVT is defined by the decomposition
\[
P(a,t,m|i,q)=P(m|i,q)\times P(a,t|m,i,q),
\]
where \(i\) is the image, \(q\) is the question, \(m\) is the mode prefix, \(t\) is the thinking process, and \(a\) is the answer [2509.22746]. This formalization separates reasoning into two coupled problems: mode selection \(P(m|i,q)\) and mode-conditioned reasoning and answering \(P(a,t|m,i,q)\). The formulation places mode choice inside the autoregressive generation process rather than treating it as a separate heuristic or an inference-time-only trick.

The need for MoVT is motivated by the complementarity of two reasoning modes. The paper distinguishes **text-based reasoning**, which reasons purely in language like standard CoT, from **visually-grounded reasoning**, which explicitly anchors reasoning to image regions, typically by emitting coordinates such as bounding boxes [2509.22746]. The argument is not that one mode should replace the other, but that no single reasoning mode dominates across all tasks. This directly motivates a model that can both learn multiple modes and differentiate among them.

A related conceptual strand comes from the “visual thoughts” perspective, which argues that multimodal chain-of-thought improves reasoning because it inserts intermediate representations that carry image information into the reasoning stream [2505.15510]. That paper states that “Visual thoughts are intermediate, logic-driven cross-modal representations that facilitate and accelerate multimodal reasoning within a unified perspective. By caching distilled visual information, they bridge raw pixels and linguistic rationales, enabling fast, context-aware access without reprocessing the image.” This suggests that MoVT is not merely a policy over output styles; it can also be understood as a policy over distinct forms of intermediate visual state.

## 2. MoVT in the visual-thought literature

The immediate background to MoVT is the attempt to unify different multimodal chain-of-thought formats. “Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought” argues that Textual-MCoT and Interleaved-MCoT are not fundamentally different reasoning paradigms; rather, both are expressions of visual thoughts [2505.15510]. It defines four visual-thought forms: **Natural Language (N-LANG)**, **Structured Language (S-LANG)**, **Edited Image (E-IMG)**, and **Generative Image (G-IMG)**. Their reported task affinities are differentiated: N-LANG is strongest for coarse perception, S-LANG is best for relation reasoning, E-IMG excels on detailed attribute reasoning, and G-IMG is best for iterative, multi-step reasoning [2505.15510].

This broader literature is relevant because MoVT can be interpreted as a restricted but operational mixture over visual-thought expression types. In the 2025 MoVT paper, the two instantiated modes are text-based reasoning and grounded reasoning [2509.22746]. In the 2025 “Visual Thoughts” paper, the space of expressions is broader and includes both textual and image-form intermediates [2505.15510]. A plausible implication is that the two-mode MoVT formulation is one concrete point within a larger design space of visual-thought mixtures.

A second related direction is “Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings,” which presents modal-mixed CoT in which the model alternates between text tokens and compact latent visual embeddings \(z_i\) [2602.00574]. There, the full generation trajectory is written as
\[
[\mathrm{BOS}]; x_{1:i}; \langle\text{START}\rangle ; z_{1:K} ;\langle\text{END}\rangle; x_{i+1:j}\; \cdots [\mathrm{EOS}],
\]
with the \(z_i\) functioning as internal visual sketches. That paper explicitly characterizes itself as a concrete realization of the broader Mixture-of-Visual-Thoughts idea. The comparison is informative: the 2025 MoVT system mixes explicit reasoning modes under prefix control, whereas the 2026 modal-mixed system interleaves language with latent visual states inside the VLM [2602.00574].

## 3. AdaVaR: supervised cold-start and RL-based mode induction

The framework used to train MoVT is **AdaVaR**: **Adaptive Visual Reasoning** [2509.22746]. AdaVaR has two stages. The first is a **supervised cold-start stage** that teaches the base model to produce multiple reasoning modes in a unified format. The second is an **RL stage** that induces context-dependent mode selection and further improves reasoning.

In the supervised stage, AdaVaR uses a uniform sequence format for all modes and assigns each mode a unique prefix token: `<ground>` for grounded reasoning and `<text>` for text-based reasoning [2509.22746]. The prompt instructs the model that grounded reasoning should output coordinates like `object[x1, y1, x2, y2]`, while text-based reasoning should reason only in text, without localization. The model response is structured as mode prefix, `<think> ... </think>`, and `<answer> ... </answer>`.

The SFT data mixes expert trajectories from both modes. The visually-grounded reasoning data is taken from high-quality SFT data from prior grounded reasoning work, especially VoCoT, and the final SFT setup reports a grounded dataset of **119K examples** [2509.22746]. The text-based reasoning data is built by distilling a text-based visual reasoning model and applying rejection sampling; in the appendix, the paper states that it distills Orsta, sampling questions from R1-OneVision, LLaVA-CoT, NuminaMath1.5, and Virgo. After rejection sampling, it reports **115K reasoning examples** and **95K direct-answer examples** [2509.22746].

A key claim is that simply mixing grounded and text-based data is not sufficient. The mode-specific prefix is described as crucial because it lets the model distinguish between reasoning modes, learn both under a unified interface, and later explore them evenly during RL [2509.22746]. This is one of the clearest ways MoVT differs from undifferentiated data mixing.

## 4. AdaGRPO and the mechanics of mode selection

The RL component of AdaVaR is **AdaGRPO**, a modification of GRPO designed for multi-mode reasoning [2509.22746]. The paper identifies two problems with vanilla GRPO in this setting: the model may under-explore modes and generate rollouts from only one mode, and rollout-level advantage alone does not explicitly compare modes.

AdaGRPO addresses this with three elements. First, it uses **prefix-guided mode exploration**: the \(2n\) rollouts are split into \(n\) text-based rollouts with \(m_t=\langle text\rangle\) and \(n\) grounded rollouts with \(m_v=\langle ground\rangle\), ensuring even exploration across modes [2509.22746]. Second, it uses a reward function following DeepSeek-R1-style rewards, with **format rewards** and **accuracy rewards**, and the accuracy reward is **1** or **0** based on rule-based evaluation of the produced answer. Third, it computes **mode-relative advantages** by modeling the rewards of the two mode-specific rollout sets as Gaussian distributions:
\[
P_t= N\left(\mu_t,\sigma_t^2\right), \qquad P_v= N\left(\mu_v,\sigma_v^2\right),
\]
and then defining
\[
A_v=\Phi\!\left(\frac{\mu_v-\mu_t}{\sqrt{\sigma_v^2+\sigma_t^2}}\right), \qquad A_t=1-A_v
\]
or equivalently
\[
A_t=\Phi\!\left(\frac{\mu_t-\mu_v}{\sqrt{\sigma_v^2+\sigma_t^2}}\right), \qquad A_v=1-A_t
\]
as reported in the appendix [2509.22746].

The token-level assignment is asymmetric by design. Mode prefix tokens receive the mode-relative advantage, while thinking tokens receive the rollout-level advantage \(A_j\) [2509.22746]. The intended effect is explicit: prefix tokens are optimized to encourage selecting the better mode, whereas reasoning tokens are optimized for the quality of the reasoning process itself.

The paper also states that RL includes a **KL penalty** with coefficient **0.04**, and it uses curriculum learning with two data phases: a **binary mixture** containing only **OmniCount** and **Geo170K**, followed by a **diverse mixture** containing the remaining RL data, including math, OCR, object counting, science, grounding, and document tasks [2509.22746]. This curriculum is intended to help the model first learn coarse mode distinctions and then refine them on harder tasks.

## 5. Empirical performance and mode-selection behavior

The reported models are **AdaVaR-3B** built from **Qwen2.5-VL-3B** and **AdaVaR-7B** built from **Qwen2.5-VL-7B** [2509.22746]. Evaluation is conducted on eight benchmarks: **MathVista**, **MathVision**, **MathVerse** (vision-only split), **WeMath**, **MMStar**, **V\***, **POPE**, and **SpatialScore**. These benchmarks cover mathematical reasoning, general perceptual reasoning, visual search, spatial reasoning, and hallucination diagnosis.

For **AdaVaR-3B**, the paper reports: MathVista **69.8**, MathVision **24.5**, MathVerse **35.2**, WeMath **33.8**, MMStar **59.3**, V\* **77.0**, POPE **88.2**, SpatialScore **18.9**, with average **50.84** [2509.22746]. For **AdaVaR-7B**, it reports: MathVista **74.4**, MathVision **28.5**, MathVerse **43.0**, WeMath **44.8**, MMStar **63.0**, V\* **83.4**, POPE **89.0**, SpatialScore **20.4**, with average **55.82** [2509.22746]. The paper highlights that AdaVaR-3B matches or approaches much larger models, that AdaVaR-7B surpasses GPT-4o in average performance, and that AdaVaR is the only model family in its table that improves across all datasets relative to the base Qwen2.5-VL models [2509.22746].

The behavioral evidence for adaptive mode selection is central. The paper reports that on **math problems**, AdaVaR tends to choose **text-based reasoning**; on **V\*** and **POPE**, it tends to choose **grounded reasoning**; and on MMStar and other mixed tasks, it makes more nuanced category-dependent choices [2509.22746]. A specific example is given for MathVista: in SFT, grounded mode is chosen **31%** of the time, whereas after RL, grounded mode choice drops to **2%**. This is presented as a clear demonstration that RL is needed for mode selection.

The ablations support the design choices. Removing both Ada-Adv and PG-Exp, which is equivalent to standard GRPO, causes the model to reinforce the SFT model’s bias toward grounded mode and substantially reduces performance [2509.22746]. Removing Ada-Adv only weakens mode selection while retaining some benefit from prefix-guided exploration. Removing diverse mixed data or removing curriculum learning also degrades performance. Most notably, removing the mode prefix yields a “Mix-SFT-RL” variant that performs worse than even single-mode baselines, which the paper treats as evidence that explicit mode identity is essential [2509.22746].

## 6. Interpretive significance, misconceptions, and limitations

One common misconception is that MoVT is merely a matter of combining heterogeneous training data. The MoVT paper explicitly argues otherwise: the model does not just “mix” data; it learns to differentiate the reasoning modes and adaptively choose among them [2509.22746]. Another misconception is that multimodal reasoning necessarily requires explicit image generation or external tools. Related work shows multiple possibilities: “Visual Thoughts” frames both textual and image-form intermediates as expressions of the same underlying mechanism [2505.15510], whereas modal-mixed CoT uses latent visual embeddings generated and consumed inside the VLM rather than explicit rendered images [2602.00574].

The mechanistic interpretation offered by the visual-thought literature is that visual thoughts serve as intermediaries between the input image and deeper transformer layers [2505.15510]. Attention analysis in that work indicates that, when visual thoughts are present, attention shifts away from the raw image and toward the visual thought; information-flow analysis suggests that most image information first flows into the visual thought and then from the visual thought into the reasoning process. This suggests that MoVT’s adaptive mode selection can be viewed not only as selecting an output style but also as selecting how visual information should be compressed and transmitted during reasoning.

The current MoVT instantiation also has explicit limitations. The paper states that AdaVaR does **not reach the upper bound** of adaptive reasoning, that mode selection is not perfect, that grounded reasoning is limited on abstract tasks, that text-based reasoning can still hallucinate or overthink, that the current system focuses on **English** scenarios, and that grounded mode struggles in cases such as multi-image reasoning and abstract concepts that cannot be grounded [2509.22746]. It suggests future extensions including more reasoning modes beyond text and grounding, a direct-answer mode, search-based or non-linear reasoning modes, explicit reasoning about which mode to use, richer RL data, and multi-round SFT+RL iteration.

Taken together, these strands support a general interpretation of MoVT: general multimodal reasoning may require a mixture of reasoning styles rather than a single canonical chain-of-thought format [2509.22746]. In the two-mode AdaVaR formulation, the mixture is between text and grounding. In the visual-thought literature, the mixture extends across N-LANG, S-LANG, E-IMG, and G-IMG [2505.15510]. In modal-mixed latent reasoning, the mixture is between language tokens and latent visual sketches [2602.00574]. The common principle is that reasoning quality depends on access to multiple representational formats and on learning when each format is appropriate.

Source: https://www.emergentmind.com/topics/mixture-of-visual-thoughts-movt