Papers
Topics
Authors
Recent
Search
2000 character limit reached

RedVTP: Efficient DVLM Inference

Updated 23 November 2025
  • RedVTP is a training-free approach that prunes unimportant visual tokens using masked-token attention to accelerate DVLM inference.
  • It computes stable importance scores after the first diffusion step to retain top-scoring tokens, significantly reducing FLOPs and latency.
  • Empirical results show substantial throughput gains and minimal accuracy loss, validating RedVTP as an efficient, retraining-free method.

RedVTP is a training-free approach for accelerating inference in diffusion vision-LLMs (DVLMs) by pruning unimportant visual tokens using masked response token attention. Operating on models such as LLaDA-V and LaViDa, RedVTP harnesses the early-stage diffusion dynamics to maximize computational efficiency while maintaining—sometimes improving—generation accuracy. The algorithm is notable for its single-shot pruning protocol, in which importance scores derived from masked-token attention are computed after the initial diffusion step, and only the top-scoring visual tokens are retained for subsequent processing steps. This process yields substantial reductions in floating point operations (FLOPs), latency, and memory requirements, and achieves state-of-the-art throughput improvements without the need for model retraining (Xu et al., 16 Nov 2025).

1. Diffusion Vision-LLM (DVLM) Inference Pipeline

DVLMs integrate visual and linguistic modalities using transformer architectures with a parallel token decoding process enabled by diffusion-based unmasking. The model architecture comprises:

  • Vision Encoder: Splits the input image into NN patches, embedding each into Rdv\mathbb{R}^{d_v}, then processing via transformer layers to yield NN visual token embeddings.
  • Projector: Maps visual token embeddings into a shared language space of dimension dd, generating the matrix V∈RN×dV \in \mathbb{R}^{N \times d}.
  • Text Encoder: Encodes an mm-token textual prompt into T∈Rm×dT \in \mathbb{R}^{m \times d}.
  • Diffusion LLM (DLM): An LL-layer, HH-head transformer with bidirectional attention, responsible for unmasking a sequence of Ï„\tau response tokens over Rdv\mathbb{R}^{d_v}0 steps.

Inference begins with a fully masked response Rdv\mathbb{R}^{d_v}1. At each step Rdv\mathbb{R}^{d_v}2,

Rdv\mathbb{R}^{d_v}3

where previously unmasked tokens remain static while masked positions are selectively unmasked based on the model's predictions. The computational complexity for a single layer on sequence length Rdv\mathbb{R}^{d_v}4 is Rdv\mathbb{R}^{d_v}5, with Rdv\mathbb{R}^{d_v}6 as the FFN hidden dimension. Since Rdv\mathbb{R}^{d_v}7, reducing Rdv\mathbb{R}^{d_v}8 (the visual token count) quadratically impacts overall FLOPs due to self-attention costs.

2. Masked-Token–Guided Visual Token Importance Scoring

RedVTP's central innovation is in computing visual token importance scores based exclusively on attention from still-masked response tokens immediately after the first diffusion step. Let Rdv\mathbb{R}^{d_v}9 be the attention matrix for head NN0 of layer NN1 at step NN2. The averaged attention over all heads and layers is:

NN3

The importance score for visual token NN4 at step NN5 is defined as:

NN6

where NN7 indexes still-masked response positions and NN8 indexes visual tokens.

Empirically, on datasets such as InfoVQA with LLaDA-V (NN9), the cosine similarity between dd0 and dd1 for dd2 exceeds 0.95, indicating that importance rankings are stable after the first step and do not benefit from recalculation in subsequent steps.

3. RedVTP Pruning Procedure

Exploiting the stability of masked-token-based importance scores, RedVTP conducts pruning once, immediately after the first diffusion step. The protocol consists of:

  • Running the initial diffusion step (dd3) on the complete set dd4 and recording all attention matrices dd5.
  • Computing the averaged attention dd6 and importance vector dd7.
  • Selecting a retention ratio dd8; the top-dd9 fraction of visual tokens by importance (V∈RN×dV \in \mathbb{R}^{N \times d}0) are retained.
  • Forming V∈RN×dV \in \mathbb{R}^{N \times d}1 by extracting the corresponding visual token rows.
  • For subsequent steps V∈RN×dV \in \mathbb{R}^{N \times d}2, progressing the diffusion process on V∈RN×dV \in \mathbb{R}^{N \times d}3.

This single-shot pruning introduces only minor algorithmic overhead while remaining entirely training-free, as no additional learning or model adaptation occurs. The algorithmic pseudocode strictly follows these steps, ensuring transparent reproducibility.

4. Computational Complexity and Efficiency

Prior to pruning, the computational cost is

V∈RN×dV \in \mathbb{R}^{N \times d}4

with V∈RN×dV \in \mathbb{R}^{N \times d}5. After pruning to V∈RN×dV \in \mathbb{R}^{N \times d}6 visual tokens, the sequence length becomes V∈RN×dV \in \mathbb{R}^{N \times d}7, and the remaining steps (V∈RN×dV \in \mathbb{R}^{N \times d}8) operate at this reduced size:

V∈RN×dV \in \mathbb{R}^{N \times d}9

Because the quadratic term mm0 dominates in large mm1 regimes, a reduction in mm2 induces significant computational savings, approaching linearity as mm3.

5. Empirical Findings

RedVTP is benchmarked on standard multimodal datasets, including Ai2D, DocVQA, RealworldQA, InfoVQA, MME, and MMBench. Summarized results for the LLaDA-V and LaViDa models are as follows:

Model Token Retention (r) Latency Reduction (%) Throughput Gain (%) Accuracy Change (%)
LLaDA-V 75% 23.11 32.35 +0.16
LLaDA-V 50% 32.04 52.75 –0.26
LLaDA-V 25% 44.57 (max 64.97) 91.66 (max 186) –4.15
LaViDa+RedVTP 75% 10.70 (max 21.87) 12.61 (max 28.05) –2.20

Notably, on InfoVQA (LLaDA-V, mm4), latency decreases from 17.31 s to 6.06 s and throughput increases from 1.79 to 5.11 tokens/s. When both LLaDA-V+RedVTP and LaViDa retain mm525% of visual tokens, RedVTP yields a +23.72% average accuracy across six benchmarks relative to LaViDa.

6. Ablation and Comparative Evaluation

Ablation studies address alternative importance scoring and pruning strategies:

  • Importance scores based solely on attention from still-masked response tokens yield the highest average accuracy at fixed retention ratios (mm6). Strategies using prompt-only, decoded-only, all-response tokens, prompt+all, visual-only, or vision-encoder-based (FoPru) scoring underperform by 1.2–6.5 accuracy points.
  • Random pruning inflicts substantial accuracy drops.
  • Progressive pruning, which performs re-scoring and token removal at each diffusion step, attains similar accuracy as RedVTP but increases latency up to 34.4% and lowers throughput by up to 25.5%, rendering it less efficient.

These ablation results demonstrate that masked-token-guided, single-step pruning constitutes the highest-utility approach for DVLM efficiency improvements.

7. Significance and Applicability

RedVTP provides an architecture-agnostic, parameter-free, and training-free mechanism for accelerating inference in diffusion-based vision-LLMs. By leveraging the empirical consistency of masked-token attention-derived visual token importance, RedVTP achieves up to 186% throughput increase and 64.97% latency reduction on canonical DVLM benchmarks, often with negligible or positive impact on accuracy. These properties enable scalable, efficient deployment of DVLMs in resource-constrained or latency-sensitive settings without retraining or complicated post-hoc adaptation (Xu et al., 16 Nov 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to RedVTP.