Papers
Topics
Authors
Recent
Search
2000 character limit reached

Value-aware Stochastic KV Cache Eviction (VaSE)

Updated 5 July 2026
  • The paper introduces a training-free recipe that protects high-magnitude value states and uses stochastic sampling to mitigate irreversible eviction errors.
  • The method combines deterministic reservation with sampling based on attention and leverage scores, ensuring critical tokens are retained for accurate reasoning.
  • Empirical evaluations on Qwen3 models reveal that VaSE improves decoding accuracy by over 4% and reduces memory usage while supporting efficient long-form generation.

Searching arXiv for the cited papers to ground the article in the current literature. arXiv search: (Chang et al., 2 Jun 2026) Value-aware Stochastic KV Cache Eviction (VaSE) is a training-free recipe for KV-cache compression in autoregressive Transformer-based reasoning that augments existing eviction heuristics in two ways: it explicitly protects large-magnitude value states and introduces randomness in the eviction decision to increase cache diversity. It was proposed to address the memory and compute bottleneck created by extended chains of thought, especially in settings where eviction-based methods had often lagged selection-based sparse attention alternatives in accuracy despite operating under a static memory budget (Chang et al., 2 Jun 2026).

1. Formal problem setting

In autoregressive Transformer-based reasoning, at decoding step tt, attention is computed over a cache of key-value pairs {(ki,vi)}i=1N\{(k_i,v_i)\}_{i=1}^N, where NN is the fixed cache budget and ki,vi∈Rdk_i,v_i\in\mathbb{R}^d. As generation proceeds, the KV cache grows; once it exceeds size NN, an eviction policy must discard some pairs to keep memory usage static. If Ct={(ki,vi)}i∈It\mathcal{C}_t=\{(k_i,v_i)\}_{i\in\mathcal{I}_t} denotes the cache before generation step tt, ∣It∣=N|\mathcal{I}_t|=N, and (knew,vnew)(k_{new},v_{new}) is the pair added after step tt, an eviction operator {(ki,vi)}i=1N\{(k_i,v_i)\}_{i=1}^N0 selects a subset of size {(ki,vi)}i=1N\{(k_i,v_i)\}_{i=1}^N1: {(ki,vi)}i=1N\{(k_i,v_i)\}_{i=1}^N2

The objective is to minimize downstream decoding loss, or equivalently to maximize accuracy, under the memory constraint. This can be written as a discrete optimization problem: at each eviction step one chooses a retained index set {(ki,vi)}i=1N\{(k_i,v_i)\}_{i=1}^N3, {(ki,vi)}i=1N\{(k_i,v_i)\}_{i=1}^N4, to maximize

{(ki,vi)}i=1N\{(k_i,v_i)\}_{i=1}^N5

Because {(ki,vi)}i=1N\{(k_i,v_i)\}_{i=1}^N6 is intractable to optimize exactly at test time, practical methods use heuristic scoring functions {(ki,vi)}i=1N\{(k_i,v_i)\}_{i=1}^N7 to rank and prune tokens. VaSE is defined as an augmentation layer over such heuristics rather than as a replacement for the underlying scorer itself (Chang et al., 2 Jun 2026).

2. Empirical motivation: value-state outliers and failure modes

The central empirical observation behind VaSE is that per-token value vectors {(ki,vi)}i=1N\{(k_i,v_i)\}_{i=1}^N8 can exhibit a heavy-tailed magnitude distribution. On Qwen3-4B over GSM8K, three magnitude proxies were measured: the {(ki,vi)}i=1N\{(k_i,v_i)\}_{i=1}^N9-norm NN0, the range NN1, and the variance NN2. All three correlate strongly, with NN3, and indicate that a small fraction, approximately NN4, of tokens have exceptionally large magnitudes (Chang et al., 2 Jun 2026).

The reported failure mode is unusually sharp. If one evicts the top-NN5 tokens ranked by NN6, model accuracy on GSM8K collapses from NN7 to NN8, and the decoder enters self-reflective loops. The paper interprets this as evidence that large-magnitude value states carry indispensable information for chain-of-thought progression. A plausible implication is that these outlier value states act as transition-critical states in long-form reasoning, so their removal perturbs the trajectory of generation more severely than their frequency in the cache would suggest (Chang et al., 2 Jun 2026).

This observation also distinguishes VaSE from heuristics based only on recent attention weights. A token may receive a low instantaneous score under an attention-based ranking while still encoding information that is disproportionately important for subsequent reasoning stability. VaSE treats this mismatch as a structural problem in eviction-based compression rather than as a mere calibration issue.

3. Stochastic retention and protected reservation

Standard eviction procedures such as SnapKV rank tokens by an importance score NN9, defined as the average attention weight over the last ki,vi∈Rdk_i,v_i\in\mathbb{R}^d0 queries, and discard the lowest ki,vi∈Rdk_i,v_i\in\mathbb{R}^d1. VaSE replaces that hard thresholding with sampling without replacement. If ki,vi∈Rdk_i,v_i\in\mathbb{R}^d2 is the set of reserved tokens and ki,vi∈Rdk_i,v_i\in\mathbb{R}^d3 is the set of remaining candidates, VaSE samples a subset ki,vi∈Rdk_i,v_i\in\mathbb{R}^d4, ki,vi∈Rdk_i,v_i\in\mathbb{R}^d5, with

ki,vi∈Rdk_i,v_i\in\mathbb{R}^d6

Under this rule, every candidate has a strictly positive retention probability at each step, ki,vi∈Rdk_i,v_i\in\mathbb{R}^d7, so no token is guaranteed to be evicted permanently. Over ki,vi∈Rdk_i,v_i\in\mathbb{R}^d8 eviction events, the cumulative keep-probability is ki,vi∈Rdk_i,v_i\in\mathbb{R}^d9, which the method presents as a mechanism for promoting long-term cache diversity (Chang et al., 2 Jun 2026).

The second ingredient is deterministic protection of large-magnitude value states. Let NN0 denote the persistent budget and NN1 the new-token buffer size. VaSE reserves NN2 slots for tokens whose value states rank among the top-NN3 by magnitude: NN4 The buffer tokens NN5 of size NN6 are always protected as well, so the reserved set is

NN7

The remaining NN8 slots are then filled by stochastic sampling from NN9 according to Ct={(ki,vi)}i∈It\mathcal{C}_t=\{(k_i,v_i)\}_{i\in\mathcal{I}_t}0. This guarantees that a large-magnitude token is not evicted simply because its current attention score is low, provided Ct={(ki,vi)}i∈It\mathcal{C}_t=\{(k_i,v_i)\}_{i\in\mathcal{I}_t}1 (Chang et al., 2 Jun 2026).

Taken together, these two mechanisms address two different pathologies of deterministic top-Ct={(ki,vi)}i∈It\mathcal{C}_t=\{(k_i,v_i)\}_{i\in\mathcal{I}_t}2 eviction: irreversible loss caused by a single below-cutoff step, and catastrophic removal of rare but unusually consequential value states.

4. Variants and algorithmic realizations

VaSE was introduced as a recipe that can be layered on top of existing eviction heuristics. The paper describes two variants (Chang et al., 2 Jun 2026):

Variant Base method Core rule
VaSE–AttnV SnapKV Reserve top-Ct={(ki,vi)}i∈It\mathcal{C}_t=\{(k_i,v_i)\}_{i\in\mathcal{I}_t}3 values by Ct={(ki,vi)}i∈It\mathcal{C}_t=\{(k_i,v_i)\}_{i\in\mathcal{I}_t}4; sample remaining tokens by Ct={(ki,vi)}i∈It\mathcal{C}_t=\{(k_i,v_i)\}_{i\in\mathcal{I}_t}5
VaSE–DKV CurDKV Reserve large-magnitude values; sample remaining tokens with probability proportional to Ct={(ki,vi)}i∈It\mathcal{C}_t=\{(k_i,v_i)\}_{i\in\mathcal{I}_t}6

In VaSE–AttnV, the key score is Ct={(ki,vi)}i∈It\mathcal{C}_t=\{(k_i,v_i)\}_{i\in\mathcal{I}_t}7 and the value score is Ct={(ki,vi)}i∈It\mathcal{C}_t=\{(k_i,v_i)\}_{i\in\mathcal{I}_t}8. The single-step pseudocode computes Ct={(ki,vi)}i∈It\mathcal{C}_t=\{(k_i,v_i)\}_{i\in\mathcal{I}_t}9 for non-buffer tokens, computes tt0, constructs tt1 as the top-tt2 indices by magnitude, forms the reserved set tt3, and samples the remaining tt4 tokens from the candidate set without replacement with tt5. The resulting cache is tt6 (Chang et al., 2 Jun 2026).

In VaSE–DKV, the method builds on CurDKV. At each eviction step tt7, it draws a fresh Gaussian projection tt8, computes leverage scores

tt9

and sets

∣It∣=N|\mathcal{I}_t|=N0

The cache then retains the reserved tokens and samples the remaining ∣It∣=N|\mathcal{I}_t|=N1 tokens from the non-reserved candidates with probability proportional to ∣It∣=N|\mathcal{I}_t|=N2. The resampling of ∣It∣=N|\mathcal{I}_t|=N3 is used to inject stochasticity into CurDKV’s otherwise deterministic ranking (Chang et al., 2 Jun 2026).

5. Evaluation, accuracy, and systems behavior

The reported evaluation uses Qwen3-4B and Qwen3-14B under Apache 2.0 on six decode-phase reasoning tasks: AIME 25/26, HMMT 25 (feb+nov), GPQA-Diamond, MATH, and LiveCodeBench-v6-Medium. For each task, the full-model average number of generated tokens ∣It∣=N|\mathcal{I}_t|=N4 is measured, the cache budget is set to ∣It∣=N|\mathcal{I}_t|=N5, corresponding to approximately ∣It∣=N|\mathcal{I}_t|=N6 of the full cache, and the buffer is fixed at ∣It∣=N|\mathcal{I}_t|=N7. Evaluation uses pass@1 accuracy averaged over 8–16 random seeds, while throughput in tokens/sec is measured on A100-80 GB with FlashAttention 2 (Chang et al., 2 Jun 2026).

At ∣It∣=N|\mathcal{I}_t|=N8 KV-cache compression on Qwen3-4B, the average pass@1 accuracies are reported as follows: Full model ∣It∣=N|\mathcal{I}_t|=N9, SeerAttention-R (knew,vnew)(k_{new},v_{new})0, SnapKV (knew,vnew)(k_{new},v_{new})1, R-KV (knew,vnew)(k_{new},v_{new})2, CurDKV (knew,vnew)(k_{new},v_{new})3, VaSE–AttnV (knew,vnew)(k_{new},v_{new})4, and VaSE–DKV (knew,vnew)(k_{new},v_{new})5. On Qwen3-14B, VaSE–DKV achieves (knew,vnew)(k_{new},v_{new})6, matching SeerAttention-R at (knew,vnew)(k_{new},v_{new})7 and exceeding R-KV at (knew,vnew)(k_{new},v_{new})8. The paper summarizes this by stating that, across six reasoning tasks, Qwen3 models using VaSE with (knew,vnew)(k_{new},v_{new})9 KV-cache compression yield higher average accuracies than the SOTA selection method at the same sparsity, while outperforming the strongest eviction method by more than tt0 (Chang et al., 2 Jun 2026).

The systems results emphasize static-memory inference. Full caching scales as tt1 in keys and values and reaches OOM beyond 32 K tokens, whereas eviction methods have static overhead approximately tt2. Throughput scales inversely with tt3; at tt4, VaSE–DKV processes 450 tok/s versus 140 tok/s for the full model on 14B. Reported peak memory is approximately 60 GB for the full model, compared with approximately 30 GB for VaSE–DKV and approximately 32 GB for VaSE–AttnV. The method is therefore positioned as a way to support FlashAttention2 while maintaining a static memory footprint for long-form reasoning (Chang et al., 2 Jun 2026).

A source of terminological ambiguity is that “VaSE” appears in two distinct senses in the 2026 literature. In "Value-Aware Stochastic KV Cache Eviction for Reasoning Models" (Chang et al., 2 Jun 2026), VaSE denotes a specific recipe: reserve the top-tt5 value-magnitude tokens and sample the rest. In "Forget Without Compromise: Nexus Sampling for Streaming KV-Cache Eviction Under Fixed Budgets" (Duong et al., 22 Jun 2026), “Value-aware stochastic eviction (VaSE)” is presented more broadly as a design prescription consisting of value estimation, stochastic retention, and budget compliance. Nexus Sampling is then described as an instantiation of that prescription through a two-term Nexus score tt6, weighted reservoir sampling, and exact fixed-budget selection (Duong et al., 22 Jun 2026).

This broader framing helps situate the original VaSE recipe among neighboring approaches. VECTOR, for example, is a deterministic plug-in for eviction-based pipelines that performs three-way token routing—retention, approximation, and eviction—using both a base importance score and an offline-calibrated reconstructability signal for values. It retains all keys, reconstructs selected values on the fly with a layerwise OLS predictor, and is explicitly non-random once tt7, tt8, and tt9 are fixed (Lin et al., 22 May 2026). Relative to VaSE, VECTOR addresses information loss through approximation rather than through stochastic retention and protected reservation.

The open problems identified for the original VaSE recipe are narrowly scoped and technically specific. The work focuses on decode-phase eviction, leaving extension to prefill-phase compression to future work; it is evaluated only on Qwen3, so application to other architectures remains to be validated; and it points to a possible interaction with quantization, noting that large-magnitude values also drive quantization error because the scaling factor {(ki,vi)}i=1N\{(k_i,v_i)\}_{i=1}^N00 becomes large, implying coarse reconstruction. The paper further identifies mechanistic interpretability of these outlier values—described as possible loop-breakers in chain-of-thought transitions—as an unresolved question (Chang et al., 2 Jun 2026).

Within this literature, VaSE is therefore best understood not merely as a cache-thinning heuristic, but as a precise response to two empirically observed weaknesses of deterministic eviction: the fragility induced by irreversible top-{(ki,vi)}i=1N\{(k_i,v_i)\}_{i=1}^N01 cutoffs and the disproportionate importance of a small set of high-magnitude value states.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Value-aware Stochastic KV Cache Eviction (VaSE).