Papers
Topics
Authors
Recent
Search
2000 character limit reached

A Simple Plug-in for Improving Eviction-Based KV Cache Compression

Published 22 May 2026 in cs.LG | (2605.23258v1)

Abstract: KV cache growth is a major bottleneck for long-context inference in LLMs. Existing methods are often dominated by binary eviction or representation approximation, which may underutilize tokens that are not critical for exact retention but are still reconstructable. We present VECTOR, a plug-and-play augmentation for eviction-based pipelines that introduces three-way token routing: retention, approximation, and eviction. VECTOR combines an importance signal from the base scorer with a reconstructability signal from an offline-calibrated regression-based value estimation. By leveraging reconstructability, VECTOR recovers useful value information that would otherwise be irreversibly lost under binary eviction, while preserving key vectors for attention routing stability. Experimental results show that VECTOR improves quality-memory trade-offs under medium-to-high compression, with especially clear gains in stricter budget regimes.

Summary

  • The paper introduces VECTOR, a plug-in that divides KV-cache tokens into exact retention, value approximation, and eviction, improving Qwen3-14B KeyDiff performance by up to 9.73 points at 90% compression.
  • The method uses offline ordinary least squares to predict values from inverse-RoPE-corrected keys, achieving layer-averaged held-out R² values from 0.6863 to 0.9392 across five model families.
  • The results show the largest gains for query-agnostic eviction methods and severe compression, while adaptive approximation budgets and stronger validation remain important for query-aware baselines.

Motivation and problem setting

Long-context inference in LLMs is bottlenecked by the linear growth of the key-value (KV) cache with sequence length. The dominant compression paradigm, importance-based eviction (e.g., SnapKV, KeyDiff, KVzip, PyramidKV), makes a binary decision per token: retain exactly or discard permanently. Under tight budgets this is destructive — evicted value information cannot be recovered. Multi-state methods such as ARKV and D2O show that non-binary allocation outperforms pure eviction, but they allocate based on token importance alone and do not model reconstructability: whether a token's representation can be accurately recovered from other cached information.

The paper under review, "A Simple Plug-in for Improving Eviction-Based KV Cache Compression" (2605.23258), proposes VECTOR (Value Estimation via Collinearity and Three-way Orthogonal Routing), a plug-and-play augmentation that converts binary retain/evict pipelines into three-way allocation: retention, approximation, and eviction. Two observations motivate the design: (i) not all important tokens are equally compressible, so reconstructability should serve as an orthogonal allocation axis; and (ii) keys and values exhibit asymmetric sensitivity to perturbation — key errors are exponentially amplified through softmax, whereas value errors propagate only linearly through the post-softmax weighted sum (2605.23258).

Method

Value estimation via K→V collinearity

VECTOR's approximation tier rests on the claim that values can be predicted from keys with a fixed linear map. Since K=WKhK = W_K h and V=WVhV = W_V h are both projections of the same hidden state hh, and since the effective rank of hh is low (the "massive activations" phenomenon; cf. MLA's shared latent), a linear predictor is feasible even when dk,dv≪dd_k, d_v \ll d. Empirically, offline OLS fitted per layer on C4 activations (10,000 sequences of 4,096 tokens) achieves held-out layer-averaged Rglobal2R^2_{\text{global}} between 0.6863 (Qwen3-30B-A3B) and 0.9392 (Qwen3-0.6B) across five model families, establishing the prerequisite for the method.

Two design choices matter here:

  • OLS over pseudo-inverse. The Moore–Penrose pseudo-inverse minimizes ℓ2\ell_2-norm solutions in hh-space, not VV-prediction error. The ablation is stark: OLS achieves positive R2R^2 on every layer of every model, while MP yields negative mean V=WVhV = W_V h0 on four of five models (catastrophically so on Qwen3-0.6B), confirming that the two objectives induce different geometries.
  • RoPE decoupling. Because RoPE applies position-dependent rotations to keys only, fitting a static OLS on post-RoPE keys would require position-dependent estimators. VECTOR applies an inexpensive inverse rotation to cached keys before reconstruction, exposing position-independent K–V collinearity.

Three-way budgeted allocation

Given target compression ratio V=WVhV = W_V h1 and approximation ratio V=WVhV = W_V h2, the pipeline proceeds in three steps. First, the base eviction scorer selects an expanded candidate pool of size V=WVhV = W_V h3. Second, per-token reconstruction error V=WVhV = W_V h4 is computed over this pool. Third, asymmetric truncation executes the budget swap: all V=WVhV = W_V h5 keys are retained exactly; the V=WVhV = W_V h6 tokens with lowest error have their values discarded and reconstructed on demand as V=WVhV = W_V h7; the remaining V=WVhV = W_V h8 tokens keep full KV pairs. The footprint matches V=WVhV = W_V h9 full pairs exactly, so memory savings are preserved while part of the information binary eviction would destroy is recovered. Approximation is applied to values only; a K-only ablation shows inconsistent gains and occasionally degrades below the unaugmented baseline at high compression (e.g., −10.95 points on HotpotQA relative to V-only at hh0).

Theoretical analysis

The paper defines an importance-weighted distortion measure hh1 in which evicted tokens contribute their full importance weight and approximated tokens contribute weight scaled by normalized reconstruction error. Proposition 1 gives a closed-form condition: expanding the approximation tier reduces distortion iff

hh2

where hh3 is the average importance in the expanded pool and hh4 the boundary score of the eviction set. The required predictability therefore depends on the skewness of the importance distribution: when evicted tokens carry negligible importance (hh5), the threshold rises and net gains become harder — which the authors later use to explain why query-aware baselines benefit less. A Gaussian-residual example further connects hh6 to the ratio hh7 of residual variance to value norm variance, via truncated-normal moments over the selected low-error tokens.

For deployment, since hh8 and hh9 vary per sample, the paper adopts the empirical rule hh0 rather than dynamic optimization — a concession the authors acknowledge as suboptimal.

Experimental results

Experiments use the KVPress framework on LongBench (16 English/code tasks) and NIAH, with Llama-3.1-8B-Instruct, Qwen3-14B, and Qwen3-0.6B, augmenting four baselines spanning query-aware (SnapKV, PyramidKV) and query-agnostic (KeyDiff, KVzip) scorers at hh1.

Baseline Model hh2 Baseline avg. +VECTOR avg. Δ
KeyDiff Qwen3-14B 0.50 40.72 47.75 +7.03
KeyDiff Qwen3-14B 0.75 32.23 41.38 +9.15
KeyDiff Qwen3-14B 0.90 22.71 32.44 +9.73
KVzip Llama-3.1-8B 0.90 41.23 45.20 +3.97

Gains concentrate in medium-to-high compression and are strongest for query-agnostic baselines, where importance and reconstructability capture largely orthogonal utility dimensions. For query-aware baselines the picture is mixed: SnapKV+VECTOR shows marginal or slightly negative effects at moderate ratios but consistent gains at hh3, while PyramidKV shows no clear improvement trend — consistent with the theoretical threshold, since these scorers already preserve high-importance tokens, leaving little headroom.

The hh4 sensitivity study confirms an inverted-U: performance rises from hh5 to a moderate value then collapses near the upper limit (at hh6, from 42.9% at hh7 to 44.3% at hh8, dropping to 41.6% at hh9). At dk,dv≪dd_k, d_v \ll d0 the peak sits at dk,dv≪dd_k, d_v \ll d1, matching the deployment formula. NIAH heatmaps at dk,dv≪dd_k, d_v \ll d2 show that VECTOR shifts failure patterns from large contiguous low-score regions to localized difficult cells, with the largest recovery on KVzip.

Limitations and open questions

The paper concedes several constraints plainly. First, dk,dv≪dd_k, d_v \ll d3 is set by a static empirical formula rather than optimized per-sample or per-layer, leaving gains unrealized where the optimum deviates. Second, benefits over query-aware baselines are modest or absent, limiting the plug-in's applicability to that family. Third, the reported results derive from single runs with a fixed seed and no significance testing, so the smaller deltas (particularly for SnapKV and PyramidKV) should be interpreted cautiously. Fourth, the theory relies on independence between reconstruction errors and importance scores, an assumption whose validity across tasks is not directly verified. Open questions include scorer-aware allocation and adaptive approximation ratios, and whether the OLS dk,dv≪dd_k, d_v \ll d4 threshold of Proposition 1 can be monitored online to trigger per-sample adjustment of dk,dv≪dd_k, d_v \ll d5.

Conclusion

VECTOR reframes eviction-based KV compression as a three-way retain–approximate–evict allocation problem, adding reconstructability — measured by offline-calibrated OLS K→V prediction error — as an allocation dimension orthogonal to token importance. The mechanism is lightweight (no retraining, no architectural change, one-time calibration), preserves exact keys for attention stability, and delivers its clearest improvements precisely where eviction methods fail most: strict budgets and query-agnostic scoring. The theory ties when the approach helps to the skewness of importance distributions and the achievable dk,dv≪dd_k, d_v \ll d6, and the empirical results are consistent with that account, including its prediction of limited headroom for query-aware baselines.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.