Token Perception: Multimodal Insights
- Token perception is defined as the process where tokens act as perceptual grounding points, influencing semantic understanding and model bias across modalities.
- Research quantifies token perceptual grounding using metrics like KL divergence, dynamic token resolution, and token reweighting in multimodal reinforcement learning.
- Innovative token operators enhance models by integrating spatial segmentation, depth encoding, and acoustic nuances to improve efficiency and diagnostic clarity.
Token perception denotes a research perspective in which tokens are treated not merely as serialization units, but as loci of perceptual grounding, perceptual control, perceptual compression, and perceptual diagnosis. Across recent work, this perspective spans language-model tokenization, region- and depth-level visual tokens, audio tokenizers, token-level reinforcement-learning credit assignment, saliency-driven token resolution, and token-centric hallucination analysis. A recurrent finding is that tokens are neither cognitively neutral nor uniformly useful: they can act as semantic primitives, as carriers of sparse visual evidence, as explicit control signals for additional perception, or as compact surrogates for high-bandwidth sensory representations (Zimmerman et al., 2024).
1. From tokenization to perceptual representation
In language modeling, tokenization is a necessary component within the current architecture of many LLMs, including transformer-based LLMs, yet its impact on model cognition is often overlooked. One line of work argues that LLMs demonstrate that the Distributional Hypothesis is sufficient for reasonably human-like language performance, while also arguing that the emergence of human-meaningful linguistic units among tokens motivates changes to linguistically agnostic tokenization techniques. In that account, tokens function both as semantic primitives and as vehicles for conveying salient distributional patterns from human language to the model; the same work further argues that tokens and pretraining can act as a backdoor for bias and other unwanted content, and that the tokenization algorithm’s objective function impacts the LLM’s cognition (Zimmerman et al., 2024).
This concern with token design extends beyond text. Deng et al.’s Perceptual Group Tokenizer is a visual backbone that entirely relies on grouping operations to extract visual features and perform self-supervised representation learning. It reported 80.3% on the ImageNet-1K self-supervised learning benchmark with linear probe evaluation, and its grouping formulation also yielded adaptive computation without re-training and explicit interpretability through grouping maps (Deng et al., 2023). In parallel, “Tokenize Anything via Prompting” defines a promptable region-level tokenizer by pairing a mask token with a semantic token for each region, jointly optimizing segmentation and concept prediction so that regional geometry and regional semantics are represented in separate but coordinated token streams (Pan et al., 2023).
Taken together, these results suggest that token perception begins before any downstream reasoning stage. The selection of token units already determines what distinctions are available, which priors are privileged, and what kinds of perceptual abstraction can be recovered later.
2. Formalizations of token perception
A central development in the literature is the replacement of informal notions of “visual grounding” or “perceptual reliance” with token-level measurements. In multimodal reinforcement learning, one widely used formulation defines token visual dependency as the Kullback-Leibler divergence between next-token distributions with and without visual conditioning:
A large indicates that visual input materially changes the predictive distribution, and thus that the token is perceptually grounded rather than primarily driven by language priors (Ye et al., 2 Apr 2026). A closely related formulation replaces the text-only comparator with a perturbed image , again using KL divergence to quantify how much the model relies on the image at generation step (Huang et al., 10 Oct 2025).
Other work refines this measurement by separating dependency from brittleness. PRPO introduces Robust Visual Dependency, combining sensitivity to strong perturbations with stability under weak perturbations:
where is the KL divergence under strong corruption and is the KL divergence under weak perturbation. Tokens with high image dependency but poor perturbation stability are gated out as brittle anchors (Li et al., 7 Jun 2026). Token-Reweighting for RLVR uses a complementary decomposition: reasoning-related tokens are identified by high prediction entropy, while perception-related tokens are identified by large log-probability changes when the image is removed (Lu et al., 26 Mar 2026).
Outside RL, Blink estimates token saliency layer by layer from attention scores and aggregates it into a patch-level saliency ratio , which governs token expansion and dropping across selected layers (Feng et al., 11 Dec 2025). In language-only settings, Marklová et al. operationalize a production-perception distinction by comparing the probabilities of the same generated tokens under production prompts and perception prompts, using a mean absolute distance between probability streams; across five open-weight models, production-perception distances exceeded production-production distances, with an overall average ratio of approximately 1.8 (Marklová et al., 13 Jul 2026).
A more abstract formulation appears in the Next Token Perception Score, which measures the overlap between autoregressive and downstream perception feature subspaces:
In the reported linear regime, NTPS both upper- and lower-bounds excess loss, correlates strongly with linear probe accuracy across 12 NLP datasets and eight pretrained models, and increases after LoRA fine-tuning in larger models (Cheng et al., 22 May 2025).
These formalizations converge on a common claim: not all tokens “see” equally, and the difference is measurable rather than metaphorical.
3. Token perception in multimodal reinforcement learning
The most concentrated use of token perception appears in multimodal RL with verifiable rewards. The basic criticism is that trajectory-level rewards assign identical learning signals across all generated tokens, even though multimodal reasoning depends on a sparse subset of visually grounded steps. “Spotlight on Token Perception for Multimodal Reinforcement Learning” reports that token visual dependency in Chain-of-Thought rollouts is sparsely distributed, with the frequency decaying exponentially, and that trajectories exhibit heterogeneous overall grounding; it then introduces Visually-Perceptive Policy Optimization, combining trajectory-level advantage shaping with token-level gradient filtering (Huang et al., 10 Oct 2025).
In that formulation, trajectories are reweighted by their average visual dependency, while gradients are applied only to the top most visually dependent tokens. On Qwen2.5-VL-7B, the reported exact-match accuracy at 8 samples rises from 55.0% with DAPO to 57.5% with VPPO; on Qwen2.5-VL-32B, performance rises from a 57.0% zero-RL baseline to 64.6% (Huang et al., 10 Oct 2025).
PGPO advances a closely related argument: only a small fraction of generated tokens genuinely depend on the image, so broadcasting one advantage to all tokens inflates gradient variance and wastes learning capacity on language-driven steps already mastered by pretraining. It introduces a threshold-gated, mass-conserving reshaping of sequence-level advantages and reports that, on Qwen2.5-VL-7B, average accuracy rises from 54.08% for a strong VPPO baseline to 54.70%, which it describes as an 18.7% relative boost over the text-only pretrained model; the extra forward pass for computing token dependency adds about 10% compute (Ye et al., 2 Apr 2026).
Token-Reweighting makes the perception-reasoning coupling itself the object of analysis. Its selective optimization study shows that reasoning-only optimization incurs an approximately 2% absolute drop versus full-token GRPO, while perception-only optimization incurs an approximately 3% drop and at low 0 can perform worse than the untrained base model. The proposed reweighting strategy then improves GRPO and DAPO across MathVerse, MathVision, MathVista, WeMath, and HalluBench, with gains of about 1–3% absolute depending on benchmark and base algorithm (Lu et al., 26 Mar 2026).
PRPO pushes the same logic to finer granularity in large vision-LLMs. It argues that existing RLVR methods assign identical learning signals to all generated tokens although only a sparse subset is causally grounded in visual evidence; the paper states that in long multimodal chains only 1 of tokens carry true visual evidence. Its Perceptual Advantage Reshaping uses RVD-derived token weights to amplify perceptually informative tokens while preserving a nonzero gradient floor for non-perceptual tokens. On seven multimodal reasoning benchmarks, PRPO reports average gains of 23.3% and 21.1% over strong baselines at 3B and 7B scales respectively, and under equal wall-clock budget it outperforms DAPO by +11.7 percentage points at 7B (Li et al., 7 Jun 2026).
A common misconception addressed by this literature is that perceptual grounding can be optimized independently of reasoning fluency, or vice versa. The RL results argue against that simplification. The evidence instead supports a sparse-and-coupled view: visually grounded tokens are few, but their interaction with reasoning tokens is structurally important.
4. Tokens as explicit perceptual operators
A second major strand treats tokens not only as diagnostic units, but as actuators that trigger or encode additional perceptual computation. “Introducing Visual Perception Token into Multimodal LLM” adds special tokens that the model generates autonomously during decoding. The Region Selection Token specifies a crop region on a 2 grid, while the Vision Re-Encoding Token uses the hidden state at a control token to guide a secondary vision encoder through cross-attention. The abstract reports that introducing these Visual Perception Tokens improves the performance of a 2B model by 23.6%, increasing its score from 0.572 to 0.708, and even outperforms a 7B parameter model by 13.4% (Yu et al., 24 Feb 2025).
Perceptio hard-wires an even more explicit spatial chain of thought. Built on InternVL, it integrates SAM2-based semantic segmentation tokens and VQ-VAE depth tokens inside the autoregressive sequence, so that the model first emits spatial tokens and then answers. To stabilize depth-token generation, it introduces composite depth-token objectives—marker, token, and count losses—together with soft merging for differentiable reconstruction. The reported gains include +0.8/+1.4/+1.1 cIoU on RefCOCO/+/g, a 10.3% improvement on HardBLINK spatial understanding accuracy, and a 1.0% increase on MMBench (Li et al., 19 Mar 2026).
“Tokenize Anything via Prompting” similarly separates “where” from “what” by adding a semantic token to each mask token in a promptable image decoder. Through joint optimization of segmentation on mask tokens and concept prediction on semantic tokens, the model is positioned as a versatile region-level image tokenizer for segmentation, open-vocabulary recognition, and captioning; the abstract reports a CIDEr score of 164.7 on the Visual Genome region captioning task with an additional 38M-parameter causal text decoder (Pan et al., 2023).
Spatial structure can also be embedded through cognitively inspired tokens. “Cognitively-Inspired Tokens Overcome Egocentric Bias in Multimodal Models” introduces perspective tokens based either on body-keypoint cues or on abstract rotation-based representations. Integrated into LLaVA-1.5-13B, these tokens improve level-2 visual perspective-taking across Isle Bricks V2, COCO, and 3DSRBench, with rotation-based tokens generalizing to non-human reference agents (Leonard et al., 23 Jan 2026).
The same principle now appears in audio. UniAudio-Token argues that semantic speech tokenizers suffer from “acoustic blindness” because ASR-oriented abstraction suppresses vocal-style nuances and non-speech events. Its Semantic-Acoustic Primitives provide supervision over linguistic content, vocal attributes, and auditory scene structure, while Semantic-Acoustic Equilibrium adaptively restores shallow acoustic detail via content-aware gating. Reported results include positive Silhouette Scores on ESC-10 and ESC-50, improved speech reconstruction with average WER reduced from 4.47 to 3.68 and MOS raised from 4.03 to 4.19, and downstream Audio-LLM understanding gains of +5–6% over the best single-codebook baselines (Song et al., 29 May 2026).
This body of work suggests that explicit perception tokens are not merely prompt engineering artifacts. They are being used as learned interfaces between symbolic generation and modality-specific perceptual operations.
5. Dynamic token resolution, merging, and communication
A third strand treats token perception as a resource-allocation problem: which tokens deserve more resolution, which can be merged, and which are worth transmitting at all. Blink is motivated by two observations in multimodal LLMs: the model’s attention “spotlight” shifts across image regions over transformer layers, and allocating more computation to high-attention tokens improves visual perception. Blink therefore combines Saliency-Guided Scanning with Dynamic Token Resolution, expanding salient tokens through a TokenSR module and dropping them when they lose focus. The paper reports 10–20% reduction in end-to-end inference FLOPs while improving or preserving performance across seven vision-language benchmarks, and on MME perception the score rises from 1505.7 for the vanilla model to 1519.7 for Blink (Feng et al., 11 Dec 2025).
MaMe and MaRe address efficiency through matrix-based token merging and restoration. MaMe is training-free, differentiable, parameter-free, and based entirely on matrix operations, avoiding sorting and scattered writes. When applied to pre-trained models, the method doubles ViT-B throughput with a 2% accuracy drop; fine-tuning the last layer with MaMe yields a 1.0% accuracy improvement at 1.1x speed; in image synthesis, the MaMe+MaRe pipeline reduces Stable Diffusion v2.1 generation latency by 31% while improving LPIPS, PSNR, CLIP Score, and FID relative to both baseline and ToMeSD (Huo et al., 15 Apr 2026).
DinoLink applies token-centric compression to bandwidth-constrained V2X perception. It combines a saliency-aware Top-3 selector with Residual Vector Quantization, transmitting codebook indices and positional priors rather than raw pixels. On nuScenes, it reports a 4 bitrate reduction while maintaining 32.8% mAP, and in a simulated LoRa setting it reduces end-to-end round-trip latency from 349.56 s to 10.11 s, a 5 acceleration (Zhu et al., 24 Jun 2026).
These methods rest on the same premise as token-level RL: token importance is highly non-uniform. The difference is operational. Instead of redistributing gradients, they redistribute compute, memory, or network bandwidth around perceptually salient tokens.
6. Hallucination, token diagnostics, and production–perception asymmetry
Token perception also functions as a diagnostic lens on generation failures. Dual-Anchor Introspective Decoding identifies, for each generated token, a Spotlight layer where visual attention and object-recognition accuracy peak and a Shadow layer where visual attention is minimal and language priors dominate. The final logits are recalibrated by adding a fraction of Spotlight logits and subtracting a fraction of Shadow logits. Reported effects include POPE accuracy improving from 81.38% to 85.08%, CHAIR6 falling from 49.6% to 35.9%, and MME perception scores increasing from 559.5 to 633.7 (Wu et al., 11 Apr 2026).
A more fine-grained result appears in “First Hallucination Tokens Are Different from Conditional Ones.” Using token-level hallucination annotations and reproduced logits on RAGTruth, the paper reports that the first hallucinated token in a span carries a stronger signal than later conditional tokens. For Llama-2-13B-chat, entropy-based AUROC for first hallucinated tokens is approximately 0.78–0.82 globally, whereas by 7 it approaches 0.50; paired bootstrap tests show the difference between first and second tokens is significant with 8 across four models (Snel et al., 28 Jul 2025). This directly contradicts the assumption that hallucination signals are homogeneous within a span.
In prompt and model comparison, Spotlight mines token patterns that distinguish systematic differences from decoding noise. It defines pattern occurrence, support, support differences, and significance filters, then presents statistically over- or under-represented token patterns to human analysts. In the controlled user study, correct identification of the target difference rose from 58% to 82% in the doctor-story condition and from 25% to 60% in the farming condition when token patterns were shown, while task time decreased and perceived difficulty fell (Hedderich et al., 22 Apr 2025).
Finally, token probabilities themselves reveal role asymmetries even in decoder-only LLMs. Marklová et al. show that re-scoring the same generated tokens under production prompts and perception-oriented prompts yields systematic divergence: production-perception distances exceeded production-production distances with non-overlapping ranges across conditions in the extended experiment, and the effect replicated across Llama-3.1-8B, EuroLLM-9B, gemma-2-9b-it, Mistral-7B-Instruct-v0.3, and Qwen2.5-7B-Instruct (Marklová et al., 13 Jul 2026).
A plausible implication is that token perception has become a unifying analysis level for alignment, evaluation, and interpretability. Rather than asking only whether a response is correct or hallucinated, this literature asks where, in token space and token time, perceptual grounding appears, disappears, or fails.