- The paper introduces VECTOR, a plug-in that divides KV-cache tokens into exact retention, value approximation, and eviction, improving Qwen3-14B KeyDiff performance by up to 9.73 points at 90% compression.
- The method uses offline ordinary least squares to predict values from inverse-RoPE-corrected keys, achieving layer-averaged held-out R² values from 0.6863 to 0.9392 across five model families.
- The results show the largest gains for query-agnostic eviction methods and severe compression, while adaptive approximation budgets and stronger validation remain important for query-aware baselines.
Motivation and problem setting
Long-context inference in LLMs is bottlenecked by the linear growth of the key-value (KV) cache with sequence length. The dominant compression paradigm, importance-based eviction (e.g., SnapKV, KeyDiff, KVzip, PyramidKV), makes a binary decision per token: retain exactly or discard permanently. Under tight budgets this is destructive — evicted value information cannot be recovered. Multi-state methods such as ARKV and D2O show that non-binary allocation outperforms pure eviction, but they allocate based on token importance alone and do not model reconstructability: whether a token's representation can be accurately recovered from other cached information.
The paper under review, "A Simple Plug-in for Improving Eviction-Based KV Cache Compression" (2605.23258), proposes VECTOR (Value Estimation via Collinearity and Three-way Orthogonal Routing), a plug-and-play augmentation that converts binary retain/evict pipelines into three-way allocation: retention, approximation, and eviction. Two observations motivate the design: (i) not all important tokens are equally compressible, so reconstructability should serve as an orthogonal allocation axis; and (ii) keys and values exhibit asymmetric sensitivity to perturbation — key errors are exponentially amplified through softmax, whereas value errors propagate only linearly through the post-softmax weighted sum (2605.23258).
Method
Value estimation via K→V collinearity
VECTOR's approximation tier rests on the claim that values can be predicted from keys with a fixed linear map. Since K=WKh and V=WVh are both projections of the same hidden state h, and since the effective rank of h is low (the "massive activations" phenomenon; cf. MLA's shared latent), a linear predictor is feasible even when dk,dv≪d. Empirically, offline OLS fitted per layer on C4 activations (10,000 sequences of 4,096 tokens) achieves held-out layer-averaged Rglobal2 between 0.6863 (Qwen3-30B-A3B) and 0.9392 (Qwen3-0.6B) across five model families, establishing the prerequisite for the method.
Two design choices matter here:
- OLS over pseudo-inverse. The Moore–Penrose pseudo-inverse minimizes ℓ2-norm solutions in h-space, not V-prediction error. The ablation is stark: OLS achieves positive R2 on every layer of every model, while MP yields negative mean V=WVh0 on four of five models (catastrophically so on Qwen3-0.6B), confirming that the two objectives induce different geometries.
- RoPE decoupling. Because RoPE applies position-dependent rotations to keys only, fitting a static OLS on post-RoPE keys would require position-dependent estimators. VECTOR applies an inexpensive inverse rotation to cached keys before reconstruction, exposing position-independent K–V collinearity.
Three-way budgeted allocation
Given target compression ratio V=WVh1 and approximation ratio V=WVh2, the pipeline proceeds in three steps. First, the base eviction scorer selects an expanded candidate pool of size V=WVh3. Second, per-token reconstruction error V=WVh4 is computed over this pool. Third, asymmetric truncation executes the budget swap: all V=WVh5 keys are retained exactly; the V=WVh6 tokens with lowest error have their values discarded and reconstructed on demand as V=WVh7; the remaining V=WVh8 tokens keep full KV pairs. The footprint matches V=WVh9 full pairs exactly, so memory savings are preserved while part of the information binary eviction would destroy is recovered. Approximation is applied to values only; a K-only ablation shows inconsistent gains and occasionally degrades below the unaugmented baseline at high compression (e.g., −10.95 points on HotpotQA relative to V-only at h0).
Theoretical analysis
The paper defines an importance-weighted distortion measure h1 in which evicted tokens contribute their full importance weight and approximated tokens contribute weight scaled by normalized reconstruction error. Proposition 1 gives a closed-form condition: expanding the approximation tier reduces distortion iff
h2
where h3 is the average importance in the expanded pool and h4 the boundary score of the eviction set. The required predictability therefore depends on the skewness of the importance distribution: when evicted tokens carry negligible importance (h5), the threshold rises and net gains become harder — which the authors later use to explain why query-aware baselines benefit less. A Gaussian-residual example further connects h6 to the ratio h7 of residual variance to value norm variance, via truncated-normal moments over the selected low-error tokens.
For deployment, since h8 and h9 vary per sample, the paper adopts the empirical rule h0 rather than dynamic optimization — a concession the authors acknowledge as suboptimal.
Experimental results
Experiments use the KVPress framework on LongBench (16 English/code tasks) and NIAH, with Llama-3.1-8B-Instruct, Qwen3-14B, and Qwen3-0.6B, augmenting four baselines spanning query-aware (SnapKV, PyramidKV) and query-agnostic (KeyDiff, KVzip) scorers at h1.
| Baseline |
Model |
h2 |
Baseline avg. |
+VECTOR avg. |
Δ |
| KeyDiff |
Qwen3-14B |
0.50 |
40.72 |
47.75 |
+7.03 |
| KeyDiff |
Qwen3-14B |
0.75 |
32.23 |
41.38 |
+9.15 |
| KeyDiff |
Qwen3-14B |
0.90 |
22.71 |
32.44 |
+9.73 |
| KVzip |
Llama-3.1-8B |
0.90 |
41.23 |
45.20 |
+3.97 |
Gains concentrate in medium-to-high compression and are strongest for query-agnostic baselines, where importance and reconstructability capture largely orthogonal utility dimensions. For query-aware baselines the picture is mixed: SnapKV+VECTOR shows marginal or slightly negative effects at moderate ratios but consistent gains at h3, while PyramidKV shows no clear improvement trend — consistent with the theoretical threshold, since these scorers already preserve high-importance tokens, leaving little headroom.
The h4 sensitivity study confirms an inverted-U: performance rises from h5 to a moderate value then collapses near the upper limit (at h6, from 42.9% at h7 to 44.3% at h8, dropping to 41.6% at h9). At dk,dv≪d0 the peak sits at dk,dv≪d1, matching the deployment formula. NIAH heatmaps at dk,dv≪d2 show that VECTOR shifts failure patterns from large contiguous low-score regions to localized difficult cells, with the largest recovery on KVzip.
Limitations and open questions
The paper concedes several constraints plainly. First, dk,dv≪d3 is set by a static empirical formula rather than optimized per-sample or per-layer, leaving gains unrealized where the optimum deviates. Second, benefits over query-aware baselines are modest or absent, limiting the plug-in's applicability to that family. Third, the reported results derive from single runs with a fixed seed and no significance testing, so the smaller deltas (particularly for SnapKV and PyramidKV) should be interpreted cautiously. Fourth, the theory relies on independence between reconstruction errors and importance scores, an assumption whose validity across tasks is not directly verified. Open questions include scorer-aware allocation and adaptive approximation ratios, and whether the OLS dk,dv≪d4 threshold of Proposition 1 can be monitored online to trigger per-sample adjustment of dk,dv≪d5.
Conclusion
VECTOR reframes eviction-based KV compression as a three-way retain–approximate–evict allocation problem, adding reconstructability — measured by offline-calibrated OLS K→V prediction error — as an allocation dimension orthogonal to token importance. The mechanism is lightweight (no retraining, no architectural change, one-time calibration), preserves exact keys for attention stability, and delivers its clearest improvements precisely where eviction methods fail most: strict budgets and query-agnostic scoring. The theory ties when the approach helps to the skewness of importance distributions and the achievable dk,dv≪d6, and the empirical results are consistent with that account, including its prediction of limited headroom for query-aware baselines.