---
title: Token Temporal Merging (TTM)
url: https://www.emergentmind.com/topics/token-temporal-merging-ttm
type: topic
---

# Token Temporal Merging (TTM)

Token Temporal Merging (TTM) is a class of algorithms, architectures, and theoretical techniques for reducing token count in sequence models by collapsing tokens deemed redundant across the temporal dimension. TTM aims to exploit the high degree of temporal correlation in video, multi-frame, time series, and related sequential transformer workloads. Across video understanding, generation, retrieval, time-series modeling, and 3D scene reconstruction, TTM achieves substantial reductions in computational and memory cost—often with minimal loss, and at times measurable gains, in qualitative and quantitative performance. State-of-the-art TTM variants rely on similarity metrics (typically cosine similarity), adaptive schedules or thresholds, spatio-temporal heuristics, or dynamic program-based segmentations to decide when and which tokens to merge, and are notable both for their training-free (plug-and-play) nature and their compatibility with existing transformer backbones.

## 1. Principles and Formulations of Token Temporal Merging

TTM generalizes the “token merging” idea (as in ToMe for images) to temporal or spatio-temporal data. For a set of embeddings $X \in \mathbb{R}^{F \times N \times D}$, where $F$ is frames (or time steps), $N$ spatial tokens, and $D$ channel dimension, TTM identifies groups of tokens across time that remain highly similar under some criterion, and merges them to a single representative. This can be formalized as matching token pairs or chains across time using

$$
\text{sim}(x_i^{(t)}, x_j^{(t+1)}) = \frac{\langle x_i^{(t)}, x_j^{(t+1)} \rangle}{\|x_i^{(t)}\| \, \|x_j^{(t+1)}\|}
$$

where merging proceeds if $\text{sim}$ exceeds a layer- or schedule-dependent threshold. In most architectures, the merged representation is the (possibly weighted) average:

$$
x_{\text{merged}} = \frac{1}{|\mathcal{C}|} \sum_{i\in\mathcal{C}} x_i,
$$

where $\mathcal{C}$ denotes tokens deemed redundant over a temporal window.

TTM can employ fixed merge budgets, adaptive thresholds (e.g., 0.7 quantile of similarities per step [2501.00946]), or dynamic-programming-based global segmentation [2505.21334]. Merging is most efficient when applied in early/fat layers (e.g., the topmost U-net blocks or pre-LLM visual token staging) but also enables inner-layer or inner-head reduction [2511.21317].

## 2. Algorithmic Variants and Implementation Details

Multiple TTM instantiations have emerged, tailored to differing sequence modeling domains:

- **Progressive Multi-Granularity (PMG)**: Alternates fine-grained spatial (frame-level) and coarse temporal (clip-level) merging in blockwise stages (TempMe [2409.01156]). Spatial merging first reduces tokens within frames, followed by cross-clip and intra-clip merges—each controlled by explicit retain ratios $(R_c, R_I)$ tuned per stage.
- **Global Redundancy-Aware Segmentation**: Uses dynamic programming over the sequence to segment frame ranges into “high similarity” chunks, maximizing prunable tokens (HoliTom [2505.21334]). Within each segment, temporal merging eliminates tokens redundant across all frames; non-redundant tokens are further subjected to spatial/cluster merging.
- **Headwise and Blocklocal TTM**: Merges tokens independently in each attention head, after reordering tokens into space-time blocks (HTTM [2511.21317]). This avoids feature collapse across heads and achieves high merge ratios at minimal cost by exploiting local temporal coherency, with complexity $O(N n_b d)$ (where $n_b$ is block size).

### Pseudocode Abstractions

Below is an archetype of one TTM layer processing step:

```python
# Partition tokens into source and destination sets (A, B)
# Compute pairwise cosine similarity S[A,B]
# Select top-k pairs by similarity exceeding threshold τ
# Merge each selected pair by averaging, drop sources
# Carry forward unmerged tokens
```

Advanced variants cache merge-pairs across denoising steps (CA-ToMe [2501.00946]), use union–find to build tracklets (STTM [2507.07990]), or merge only within causal windows for time-series decoders [2405.17951].

## 3. Integration into Model Architectures

TTM is highly modular and appears in diverse positions within modern deep learning architectures:

- **Video Transformers**: TTM can be plugged either after the patch embedding or after each transformer block. In joint space-time models (ViViT, VideoMAE), tokens from all frames are eligible for merging in a global stage, whereas divided models (TimeSformer) perform TTM per frame for compatibility with spatial and temporal attention factoring [2506.03885].
- **Diffusion Models**: CA-ToMe applies TTM before MHSA in the topmost encoder/decoder UNet blocks, using adaptive similarity thresholds and caching merge-pairs to minimize redundant computation (Stable Diffusion v1.5: 1.24× speedup, minimal FID change [2501.00946]). VidToMe interleaves chunkwise local and global token alignment, merging, and delayed unmerging for efficient video editing with strong temporal consistency [2312.10656].
- **Video LLMs**: HoliTom and STTM both integrate outer-LLM TTM for aggressive pre-LLM token downsampling, followed by finer inner-LLM merging [2505.21334, 2507.07990]. KV-cache reuse is possible due to TTM’s query-agnostic nature.
- **Time Series and SSMs**: Local merging performs TTM over a bounded window of size $k$, enforcing causal constraints in decoder blocks (no future-to-past merges), and interpolating between quadratic and linear cost [2405.17951].

## 4. Computational and Practical Benefits

TTM addresses the quadratic scaling of self-attention and broad memory bottlenecks:

| Model/System                | Token Reduction | Throughput Gain | Accuracy Loss |
|----------------------------|-----------------|-----------------|--------------|
| TempMe (ViT-B/16, 12 fr)   | 95%             | 1.8×            | +5.3 R-Sum   |
| CA-ToMe (diffusion)        | ~30%            | 1.24×           | ΔFID <0.4    |
| Video-TTM, ViViT           | ~60%            | 2.5×            | <1%           |
| HoliTom (video LLM)        | 93% (prefill)   | 2.3× TTFT, 1.3× decode | 0.9%   |
| VidToMe (video editing)    | ~60% (mem.)     | ~3×             | -- (improved consistency) |

TTM generally yields $1.3$–$4\times$ speedup on video transformers, $1.2$–$7\times$ on 3D scene reconstruction, up to $54\times$ on large time-series foundation models, and $2$–$3\times$ memory savings in multi-frame diffusion models. Fine-tuning or retraining is often not required; in some cases, inference-only merging acts as a low-pass filter in time series, actually improving prediction error metrics when high-frequency noise dominates [2405.17951].

## 5. Design Choices, Schedules, and Limitations

TTM exposes several key degrees of freedom:

- **Merge Budget and Scheduling**: Merge ratio $r$ per layer governs the speed–accuracy tradeoff: constant, increasing (late), and decreasing (early) schedules yield varying loss profiles, with increasing schedules being safest for fine-grained tasks [2506.03885].
- **Similarity Thresholding**: Adaptive, per-step or global similarity thresholds (e.g., $\tau_t \approx 0.7$ in CA-ToMe) control the aggressiveness of merging [2501.00946].
- **Block and Head Granularity**: Merging can be performed at the block-level in spatio-temporal space, and independently per head to preserve representational diversity [2511.21317].
- **Window and Causality Constraints**: In time-series, merging is typically restricted to local neighborhoods and causal orders to prevent violation of temporal dependencies [2405.17951].

Limitations include possible loss of detail in highly dynamic content when merging is too aggressive, difficulty handling arbitrary sequence lengths or scene changes in fixed-schedule frameworks, and—except in architectures like STTM or HoliTom—no built-in mechanism for global, content-sensitive scheduling [2505.21334, 2507.07990]. TTM in decoders remains a less-explored area, with causal merging mandatory to preserve autoregressive correctness [2405.17951].

## 6. Empirical Evaluation and Comparisons

Comprehensive benchmarking of TTM variants demonstrates systematic advantages over hard-dropping or random replacement baselines. In video understanding and action recognition, attention-based TTM outperforms token dropping by ~4–5% accuracy for the same reduction [2506.03885]. TempMe improves GFLOPS and memory more than ToMe at equivalent accuracy, and HoliTom achieves near-lossless performance (99.1% retention) at <7% FLOPs [2409.01156, 2505.21334]. VidToMe establishes that TTM is critical for eliminating “flicker” and ensuring cross-frame consistency in diffusion-based video editing, as ablation of TTM steps sharply degrades interpolation and perceptual metrics [2312.10656]. In time series, spectral analysis reveals classes of problems (with redundant or low-rank dynamics) for which moderate token merging can even improve downstream error [2405.17951].

## 7. Extensions and Future Directions

Emerging research seeks to (1) further generalize TTM with content-aware or adaptive length scheduling [2409.01156], (2) fuse or replace averaging-based merges with more flexible, learned gating for potentially superior information retention, (3) exploit per-head and per-block adaptivity for multi-modal or scene-graph models [2511.21317], (4) extend to tasks such as video question answering and captioning (TempMe, STTM), and (5) analyze TTM’s relation to low-pass filtering and data compressibility via spectral signatures [2405.17951]. A plausible implication is that further harmonizing spectral analysis with dynamic merging schedules may enable even greater efficiency with minimal accuracy loss, particularly in streaming or real-time deployment settings.

---

**References**
- [2501.00946]: Cached Adaptive Token Merging: Dynamic Token Reduction and Redundant Computation Elimination in Diffusion Model
- [2505.21334]: HoliTom: Holistic Token Merging for Fast Video Large Language Models
- [2409.01156]: TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
- [2506.03885]: Video, How Do Your Tokens Merge?
- [2511.21317]: HTTM: Head-wise Temporal Token Merging for Faster VGGT
- [2405.17951]: Efficient Time Series Processing for Transformers and State-Space Models through Token Merging
- [2507.07990]: Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
- [2312.10656]: VidToMe: Video Token Merging for Zero-Shot Video Editing

Source: https://www.emergentmind.com/topics/token-temporal-merging-ttm