HoloV: Holistic Visual Token Pruning
- HoloV is a training-free, plug-and-play framework for multimodal large language models that prunes visual tokens holistically via adaptive crop-wise allocation.
- It combines intra-crop token diversity with attention cues to prevent representational collapse and ensure spatially distributed, semantically diverse token retention.
- Empirical results across LLaVA-1.5, LLaVA-NeXT, and Video-LLaVA demonstrate robust accuracy retention and significant latency reductions under aggressive pruning ratios.
HoloV is a training-free, plug-and-play visual token pruning framework for efficient inference in multimodal LLMs (MLLMs). It was introduced to address the computational overhead induced by massive visual token sequences, and it redefines token retention from a holistic perspective rather than preserving only locally “highlighted” tokens selected by text-vision cross-attention or attention. Its central mechanism is adaptive crop-wise allocation of the pruning budget, designed to preserve global visual context, semantic diversity, and spatial coverage under aggressive pruning ratios; the reported evaluations span image, video, and high-resolution settings across LLaVA-1.5, LLaVA-NeXT, Qwen2.5-VL-7B, and Video-LLaVA (Zou et al., 3 Oct 2025).
1. Problem formulation and motivation
HoloV was proposed in response to a specific limitation of recent efficient-inference methods for MLLMs: although token pruning reduces compute, attention-first pruning approaches tend to preserve semantically similar tokens and therefore incur pronounced performance drops under high pruning ratios. In the formulation reported for HoloV, prior methods such as FastV, SparseVLM, and FasterVLM typically use text-vision cross-attention or attention to rank visual tokens, after which top- tokens are retained and the remainder discarded (Zou et al., 3 Oct 2025).
The core critique is that high-attention tokens are frequently clustered within salient local regions or affected by positional bias. Under aggressive pruning, this produces redundancy rather than coverage: retained tokens may represent only isolated salient features, while task-relevant background, spatial relationships, or semantically distinct regions are removed. HoloV is explicitly designed to mitigate this “representational collapse” by distributing token retention across spatial crops and combining saliency with intra-crop semantic variance.
This framing places HoloV within efficient MLLM inference rather than model pretraining. The method is described as model-agnostic and compatible with multiple architectures, vision encoders, and both image and video tasks. A plausible implication is that HoloV targets the inference bottleneck while avoiding retraining costs that would otherwise complicate deployment.
2. Crop-wise holistic pruning mechanism
The HoloV pipeline begins by partitioning visual tokens from the vision encoder into spatial crops. For crop , the visual embeddings are denoted . Within each crop, HoloV computes an intra-crop similarity matrix
and defines the token diversity score as
This diversity term is combined with the crop-level attention score to form the holistic token score
where 0 is an adaptive scaling factor (Zou et al., 3 Oct 2025).
The pruning budget is then distributed across crops through adaptive quota allocation. Crop importance is defined as
1
with 2 controlling allocation sharpness. Each crop receives a quota
3
where 4 is the total number of tokens to retain. Quotas are then balanced via iterative reallocation so that the token budget is utilized without over-concentration in a few crops. Within each crop, the retained set is obtained by selecting the top-5 tokens according to 6, with 7.
An optional extension, termed fast visual context refetching, is used when inference uncertainty is high. In that case, pruned tokens are refetched through a simple FFN-based mechanism to supplement missing context without reintroducing major compute overhead. This suggests a two-stage pruning regime in which the base selection mechanism is lightweight, while limited recovery is reserved for uncertain cases.
3. Architectural intent and theoretical framing
HoloV’s design objective is not merely to retain “important” tokens, but to preserve holistic visual context. The paper characterizes this through three linked ideas: spatial spread, semantic diversity, and adaptive allocation. By combining attention saliency with intra-crop variance, the method aims to avoid selecting clusters of near-duplicate tokens and instead retain tokens that represent scattered key objects, relationships, and backgrounds. In the reported interpretation, this is what prevents the collapse of semantic representation under high pruning ratios (Zou et al., 3 Oct 2025).
The theoretical section states two formal properties. First, under an assumption of local semantic stability, HoloV’s pruning strategy is shown to bound semantic distortion in transformer outputs as a function of diversity and attention thresholds. Second, the crop-wise allocation procedure can be modeled as maximizing a monotone submodular function, with proven near-optimality under greedy selection. These claims position HoloV as more than a heuristic ranking rule: it is presented as a structured approximation to context-preserving subset selection.
A plausible interpretation is that HoloV replaces globally competitive token ranking with a constrained allocation process in which different spatial regions compete for local quotas rather than a single global top-8. That distinction is central to the framework’s reported robustness at aggressive pruning ratios.
4. Empirical evaluation across image, video, and high-resolution regimes
The reported experiments cover image-task benchmarks including GQA, MMBench, MMBench-CN, MME, POPE, SQA, VQA-v2, TextVQA, MM-Vet, and VizWiz, as well as video QA benchmarks MSVD-QA and MSRVTT-QA. The tested architectures are LLaVA-1.5 (7B, 576 tokens base), LLaVA-NeXT (7B, up to 2880 tokens), Video-LLaVA, and Qwen2.5-VL-7B. The main pruning ratios are 9, 0, and 1, with extreme pruning up to 2 (Zou et al., 3 Oct 2025).
For LLaVA-1.5, the headline reported accuracy-retention numbers are 3 at 192 retained tokens, 4 at 128 retained tokens, and 5 at 64 retained tokens, corresponding respectively to pruning 6, 7, and 8 of the original 576-token budget. For LLaVA-NeXT, at 9 pruning with 320 retained tokens from a 2880-token baseline, HoloV reports 0 retained performance. At 1 pruning, the paper reports up to 2 reduction in latency and 3 reduction in inference time, with only a 4 drop in accuracy.
| Setting | Pruning / retained tokens | Reported retained performance |
|---|---|---|
| LLaVA-1.5 | 5 pruning / 192 tokens | 6 |
| LLaVA-1.5 | 7 pruning / 128 tokens | 8 |
| LLaVA-1.5 | 9 pruning / 64 tokens | 0 |
| LLaVA-NeXT | 1 pruning / 320 tokens | 2 |
The video results are also reported as favorable. At 3 pruning, HoloV matches or outperforms prior methods with 4 accuracy on MSVD-QA and MSRVTT-QA. Cross-architecture generalization is emphasized by the Qwen2.5-VL-7B results: at 5 pruning, HoloV reports 6 average retention versus 7 for FastV. On POPE and MME, which are used as fine-grained and hallucination-oriented benchmarks, HoloV is described as substantially more robust than competing methods at high pruning.
5. Comparison with attention-first pruning and reported ablations
The main comparison class for HoloV is attention-first pruning. In the paper’s summary, methods such as FastV, SparseVLM, and generic 8-pruning are characterized as local-attention-based, non-adaptive in allocation, and vulnerable to severe performance drops at high pruning ratios. By contrast, HoloV is described as preserving context through crop-wise, diversity-aware token scoring and per-crop quota allocation (Zou et al., 3 Oct 2025).
The qualitative analysis states that HoloV retains spatially distributed, semantically distinct tokens, whereas FastV and similar methods cluster their selections in local high-attention regions. This is presented as the empirical manifestation of the framework’s holistic context retention principle. The paper also reports that varying the number of crops, or using adaptive crop numbers, does not degrade performance, which is used to support the method’s robustness.
A frequent misconception in this literature is that aggressive pruning fails only because too few tokens are kept. HoloV argues for a different diagnosis: failure arises because retained tokens are redundant and spatially concentrated. Another misconception is that saliency and coverage are interchangeable. The reported results suggest that saliency-only ranking is insufficient once pruning becomes aggressive, particularly in high-resolution and multi-frame settings.
6. Scope of the name and potential ambiguity
Within the MLLM literature, HoloV refers specifically to the visual token pruning framework described above (Zou et al., 3 Oct 2025). However, the string “HoloV” appears in other contexts, and the term is therefore context-sensitive.
In the HoloVIC dataset paper, the authors explicitly state that there is no evidence of a separate dataset named “HoloV”; all mentions in that context refer to “HoloVIC,” the large-scale multi-sensor holographic vehicle-infrastructure cooperation dataset (Ma et al., 2024). By contrast, the later paper on dense holographic associative memories uses “HoloV” in a different sense, describing it as “Holographic Volume” memory and positioning it as an optical memory primitive rather than an MLLM inference method (Brady et al., 16 Jun 2026). A separate holographic display paper uses “HoloV-style” only as a forward-looking reference to immersive display experiences, not as the name of the pruning framework (Li et al., 27 Nov 2025).
This suggests that “HoloV” is not a universally unique designation across the broader literature. In current usage, its most clearly defined technical meaning is the MLLM token-pruning method that emphasizes holistic visual context retention under aggressive pruning (Zou et al., 3 Oct 2025).