Papers
Topics
Authors
Recent
Search
2000 character limit reached

X-Slim: Cache Accelerator for Diffusion Models

Updated 21 December 2025
  • X-Slim is a unified, training-free accelerator that leverages caching of redundant computations across time, structure, and space in diffusion models.
  • It uses a dual-threshold controller to skip entire denoising steps and selectively refresh blocks and tokens while maintaining high visual fidelity.
  • Empirical evaluations demonstrate speedups up to 4.97× with negligible quality loss across tasks, establishing a new Pareto frontier in diffusion acceleration.

X-Slim (eXtreme-Slimming Caching) is a unified, training-free, cache-based accelerator designed for diffusion model inference, which systematically exploits redundant computation across temporal, structural, and spatial axes. By dynamically controlling cache usage via context-aware indicators and a dual-threshold strategy, X-Slim achieves substantial speedups for generative models with negligible perceptual loss, effectively advancing the achievable speed–quality frontier in large-scale diffusion-based synthesis tasks (Wen et al., 14 Dec 2025).

1. Motivation and Background

Diffusion models involve iterative denoising over TT timesteps, each executing a deep stack of LL transformer or U-Net blocks on NN input tokens. Inference computation and latency scale as O(T⋅L⋅N)\mathcal{O}(T \cdot L \cdot N). Primary cost drivers include: (a) temporal steps (T∼50T \sim 50 in contemporary settings); (b) structural depth (LL up to hundreds for high-fidelity tasks); and (c) spatial extent (N≫1N \gg 1k for high-resolution synthesis), with spatial FLOPs increasing quadratically due to self-attention.

Prior acceleration approaches target singular axes, such as step-level skipping (aggressive, prone to quality loss), block-level selection (safer, less savings), or token-level refreshing. These realize only local optima, leaving significant redundancy untapped.

X-Slim’s key premise is that diffusion model features exhibit strong similarity, not only across adjacent timesteps but also within certain blocks and for spatial tokens (especially background regions). Exploiting these redundancies multidimensionally enables more aggressive acceleration while minimizing the risk of error accumulation and perceptual degradation (Wen et al., 14 Dec 2025).

2. Core Architecture and Mechanisms

X-Slim is structured around a dual-threshold controller implementing a “push-then-polish” principle:

  • Reuse is “pushed” aggressively at the timestep level until error reaches an early-warning boundary.
  • Thereafter, lightweight block- and token-level refresh policies “polish” remaining accumulation.
  • Upon crossing the critical error threshold, full inference is triggered to refresh all cached features, resetting error.

At each step tt, cumulative reuse error is tracked via:

E(t)=∑k=tcalc+1t∥Δk−Δk−1∥1∥Δk−1∥1E(t) = \sum_{k = t_{\mathrm{calc}} + 1}^{t} \frac{\lVert \Delta_k - \Delta_{k-1} \rVert_1}{\lVert \Delta_{k-1} \rVert_1}

where Δk=Ok−Ik\Delta_k = O_k - I_k (true feature change at step LL0), and LL1 denotes the last step with full computation. Thresholds LL2 (“early-warning”) and LL3 (“critical”) govern the transition between coarse skipping, partial refresh, and full inference.

Temporal Level: Skipping entire denoising steps when LL4 by reusing LL5.

Structural (Block) Level: For block LL6, compute per-block relative change,

LL7

and accumulate LL8 over skipped blocks. Block-level refresh is triggered when LL9.

Spatial (Token) Level: For token NN0, compute

NN1

and only recompute the subset of tokens (top-NN2 by change) exceeding a threshold.

3. Mathematical Description and Threshold Calibration

Thresholds are set as NN3 and NN4, where NN5 corresponds to the plateau of the U-curve for NN6 (empirically calibrated per model/task) and NN7 is the maximal tolerable error (see ablation in supplementary materials).

The relative ratio NN8 is tuned, empirically NN9 yields stable speed–quality trade-off. Error is accumulated until O(T⋅L⋅N)\mathcal{O}(T \cdot L \cdot N)0, at which point full inference is run and the cache is reset.

Expected per-step cost is given by:

O(T⋅L⋅N)\mathcal{O}(T \cdot L \cdot N)1

where O(T⋅L⋅N)\mathcal{O}(T \cdot L \cdot N)2, O(T⋅L⋅N)\mathcal{O}(T \cdot L \cdot N)3, O(T⋅L⋅N)\mathcal{O}(T \cdot L \cdot N)4 are respective reuse ratios and O(T⋅L⋅N)\mathcal{O}(T \cdot L \cdot N)5, O(T⋅L⋅N)\mathcal{O}(T \cdot L \cdot N)6 are per-unit costs for block and token refreshes. Speedup is O(T⋅L⋅N)\mathcal{O}(T \cdot L \cdot N)7.

4. Algorithmic Realization

At each timestep:

  1. Accumulate error O(T⋅L⋅N)\mathcal{O}(T \cdot L \cdot N)8.
  2. If O(T⋅L⋅N)\mathcal{O}(T \cdot L \cdot N)9 (skip): T∼50T \sim 500.
  3. If T∼50T \sim 501 (full inference): T∼50T \sim 502; reset all deltas and set T∼50T \sim 503.
  4. Else (refresh): recompute select blocks (T∼50T \sim 504) and top-T∼50T \sim 505 high-change tokens; reuse cache for all other units.

Below is a workflow summary:

Mode Trigger Condition Action
Skip Step T∼50T \sim 506 Step-level cache reuse
Refresh Step T∼50T \sim 507 Recompute selected blocks/tokens
Full Inference T∼50T \sim 508 Run all layers, reset error/caches

5. Empirical Results and Comparative Performance

Empirical evaluation across diverse generators:

  • On FLUX.1-dev (T=50), X-Slim(C2F)-fast attains a 4.97T∼50T \sim 509 speedup at ImageRewardLL00.9806 (ΔLL1–0.0080), surpassing TeaCache (3.24LL2) and TaylorSeer (2.35LL3).
  • For HunyuanVideo (video, 50 steps/81 frames), achieves 3.52LL4 acceleration, VBench score of 81.69%, LPIPSLL50.1638.
  • DiT-XL/2 (ImageNet classLL6image): 3.13LL7 acceleration, FIDLL82.42, sFIDLL94.59.

Observed effects:

  • X-Slim strictly dominates alternative caching methods across the latency–quality curve, establishing a new Pareto frontier.
  • Failure manifests as structural artifacts or blurring when full refresh is too infrequent, or when aggressive reuse is employed outside the robust error corridor (before N≫1N \gg 10 or after N≫1N \gg 11).

6. Integration Guidelines and Usage Considerations

Key steps for effective deployment:

  • Calibrate thresholds using U-curve analysis of relative feature changes on a small validation set.
  • Set N≫1N \gg 12 for full computation every 8–15 skipped steps; then select N≫1N \gg 13 with N≫1N \gg 14.
  • During inference, insert the dual-threshold controller into the sampling loop; maintain caches at block and token granularity.
  • For refresh steps, set N≫1N \gg 15 and token refresh ratio to target N≫1N \gg 1680% cache reuse.

The X-Slim procedure is plug-and-play and requires no model retraining or fine-tuning.

7. Context and Significance

X-Slim introduces the first framework to exploit cacheable redundancy jointly across time, structure, and space in diffusion models. By combining aggressive timestep skipping with targeted structural and spatial refreshes, and by providing practical, empirical guidance for parameter selection, X-Slim delivers up to 4.97N≫1N \gg 17 acceleration for image generation, 3.52N≫1N \gg 18 for video, and 3.13N≫1N \gg 19 in classtt0image tasks, all while maintaining high perceptual fidelity (Wen et al., 14 Dec 2025). This approach represents a rigorous advancement of the diffusion acceleration landscape, with broad applicability to transformer-based generative architectures.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to X-Slim (eXtreme-Slimming Caching).