---
title: Value-aware Stochastic KV Cache Eviction (VaSE)
url: https://www.emergentmind.com/topics/value-aware-stochastic-kv-cache-eviction-vase
type: topic
---

# Value-aware Stochastic KV Cache Eviction (VaSE)

Searching arXiv for the cited papers to ground the article in the current literature.
arXiv search: 2606.03928
Value-aware Stochastic KV Cache Eviction (VaSE) is a training-free recipe for KV-cache compression in autoregressive Transformer-based reasoning that augments existing eviction heuristics in two ways: it explicitly protects large-magnitude value states and introduces randomness in the eviction decision to increase cache diversity. It was proposed to address the memory and compute bottleneck created by extended chains of thought, especially in settings where eviction-based methods had often lagged selection-based sparse attention alternatives in accuracy despite operating under a static memory budget [2606.03928].

## 1. Formal problem setting

In autoregressive Transformer-based reasoning, at decoding step \(t\), attention is computed over a cache of key-value pairs \(\{(k_i,v_i)\}_{i=1}^N\), where \(N\) is the fixed cache budget and \(k_i,v_i\in\mathbb{R}^d\). As generation proceeds, the KV cache grows; once it exceeds size \(N\), an eviction policy must discard some pairs to keep memory usage static. If \(\mathcal{C}_t=\{(k_i,v_i)\}_{i\in\mathcal{I}_t}\) denotes the cache before generation step \(t\), \(|\mathcal{I}_t|=N\), and \((k_{new},v_{new})\) is the pair added after step \(t\), an eviction operator \(\mathcal{E}\) selects a subset of size \(N\):
\[
\mathcal{E}(\mathcal{C}_t\cup\{(k_{new},v_{new})\})\subset\mathcal{C}_t\cup\{(k_{new},v_{new})\}, \qquad |\mathcal{E}(\cdot)|=N.
\]

The objective is to minimize downstream decoding loss, or equivalently to maximize accuracy, under the memory constraint. This can be written as a discrete optimization problem: at each eviction step one chooses a retained index set \(R\subseteq\{1,\dots,N+1\}\), \(|R|=N\), to maximize
\[
\mathcal{L}\bigl(\{(k_i,v_i)\}_{i\in R}\bigr)
=
\mathbb{E}_{\text{prompts}}\bigl[\text{accuracy or log-prob under cache }R\bigr].
\]
Because \(\mathcal{L}\) is intractable to optimize exactly at test time, practical methods use heuristic scoring functions \(s_i\) to rank and prune tokens. VaSE is defined as an augmentation layer over such heuristics rather than as a replacement for the underlying scorer itself [2606.03928].

## 2. Empirical motivation: value-state outliers and failure modes

The central empirical observation behind VaSE is that per-token value vectors \(v_i\in\mathbb{R}^d\) can exhibit a heavy-tailed magnitude distribution. On Qwen3-4B over GSM8K, three magnitude proxies were measured: the \(L_2\)-norm \(\|v_i\|_2\), the range \(\mathrm{Range}(v_i)=\max_j v_{i,j}-\min_j v_{i,j}\), and the variance \(\tfrac{1}{d}\sum_j(v_{i,j}-\bar v_i)^2\). All three correlate strongly, with \(r>0.8\), and indicate that a small fraction, approximately \(1\%\), of tokens have exceptionally large magnitudes [2606.03928].

The reported failure mode is unusually sharp. If one evicts the top-\(B\) tokens ranked by \(\mathrm{Range}(v_i)\), model accuracy on GSM8K collapses from \(80\%\) to \(14\%\), and the decoder enters self-reflective loops. The paper interprets this as evidence that large-magnitude value states carry indispensable information for chain-of-thought progression. A plausible implication is that these outlier value states act as transition-critical states in long-form reasoning, so their removal perturbs the trajectory of generation more severely than their frequency in the cache would suggest [2606.03928].

This observation also distinguishes VaSE from heuristics based only on recent attention weights. A token may receive a low instantaneous score under an attention-based ranking while still encoding information that is disproportionately important for subsequent reasoning stability. VaSE treats this mismatch as a structural problem in eviction-based compression rather than as a mere calibration issue.

## 3. Stochastic retention and protected reservation

Standard eviction procedures such as SnapKV rank tokens by an importance score \(\bar\alpha_i\), defined as the average attention weight over the last \(B\) queries, and discard the lowest \(B\). VaSE replaces that hard thresholding with sampling without replacement. If \(\mathcal{R}\) is the set of reserved tokens and \(\mathcal{C}\) is the set of remaining candidates, VaSE samples a subset \(S\subset\mathcal{C}\), \(|S|=N-|\mathcal{R}|\), with
\[
\Pr(i\in S)\propto\bar\alpha_i\quad\forall i\in\mathcal{C},
\qquad
\Pr(S)=\frac{\prod_{i\in S}\bar\alpha_i}{\sum_{T:\,|T|=N-|\mathcal R|}\prod_{j\in T}\bar\alpha_j}.
\]
Under this rule, every candidate has a strictly positive retention probability at each step, \(\pi_i^{(t)}>0\), so no token is guaranteed to be evicted permanently. Over \(T\) eviction events, the cumulative keep-probability is \(\prod_{t=1}^T\pi_i^{(t)}>0\), which the method presents as a mechanism for promoting long-term cache diversity [2606.03928].

The second ingredient is deterministic protection of large-magnitude value states. Let \(K\) denote the persistent budget and \(B\) the new-token buffer size. VaSE reserves \(v<K\) slots for tokens whose value states rank among the top-\(v\) by magnitude:
\[
\mathcal{R}_v=\mathsf{Top}_{v}\bigl\{\mathrm{Range}(v_i)\colon i\in\mathcal{I}\bigr\},
\qquad
|\mathcal{R}_v|=v.
\]
The buffer tokens \(\mathcal{B}\) of size \(B\) are always protected as well, so the reserved set is
\[
\mathcal{R}=\mathcal{R}_v\cup\mathcal{B}.
\]
The remaining \(K-|\mathcal{R}|\) slots are then filled by stochastic sampling from \(\mathcal{I}\setminus\mathcal{R}\) according to \(\bar\alpha_i\). This guarantees that a large-magnitude token is not evicted simply because its current attention score is low, provided \(|\mathcal{R}_v|<K\) [2606.03928].

Taken together, these two mechanisms address two different pathologies of deterministic top-\(K\) eviction: irreversible loss caused by a single below-cutoff step, and catastrophic removal of rare but unusually consequential value states.

## 4. Variants and algorithmic realizations

VaSE was introduced as a recipe that can be layered on top of existing eviction heuristics. The paper describes two variants [2606.03928]:

| Variant | Base method | Core rule |
|---|---|---|
| VaSE–AttnV | SnapKV | Reserve top-\(v\) values by \(\mathrm{Range}(v_i)\); sample remaining tokens by \(\bar\alpha_i\) |
| VaSE–DKV | CurDKV | Reserve large-magnitude values; sample remaining tokens with probability proportional to \(s_i=\ell_i^{(K)}\cdot\ell_i^{(V)}\) |

In VaSE–AttnV, the key score is \(\bar\alpha_i\) and the value score is \(\mathrm{Range}(v_i)\). The single-step pseudocode computes \(\bar\alpha_i\) for non-buffer tokens, computes \(\mathrm{mag}_i=\mathrm{Range}(v_i)\), constructs \(R_{\text{val}}\) as the top-\(v\) indices by magnitude, forms the reserved set \(R=R_{\text{val}}\cup B_{\text{ids}}\), and samples the remaining \(m=K-|R|\) tokens from the candidate set without replacement with \(\Pr(i\in S)\propto\bar\alpha_i\). The resulting cache is \(R\cup S\) [2606.03928].

In VaSE–DKV, the method builds on CurDKV. At each eviction step \(t\), it draws a fresh Gaussian projection \(G_t\in\mathbb{R}^{d\times r}\), computes leverage scores
\[
\ell_i^{(K)}=\|G_t^\top k_i\|_2^2,\qquad
\ell_i^{(V)}=\|G_t^\top v_i\|_2^2,
\]
and sets
\[
s_i=\ell_i^{(K)}\cdot\ell_i^{(V)}.
\]
The cache then retains the reserved tokens and samples the remaining \(N-|\mathcal{R}|\) tokens from the non-reserved candidates with probability proportional to \(s_i\). The resampling of \(G_t\) is used to inject stochasticity into CurDKV’s otherwise deterministic ranking [2606.03928].

## 5. Evaluation, accuracy, and systems behavior

The reported evaluation uses Qwen3-4B and Qwen3-14B under Apache 2.0 on six decode-phase reasoning tasks: AIME 25/26, HMMT 25 (feb+nov), GPQA-Diamond, MATH, and LiveCodeBench-v6-Medium. For each task, the full-model average number of generated tokens \(T\) is measured, the cache budget is set to \(K\approx T/4\), corresponding to approximately \(25\%\) of the full cache, and the buffer is fixed at \(B=64\). Evaluation uses pass@1 accuracy averaged over 8–16 random seeds, while throughput in tokens/sec is measured on A100-80 GB with FlashAttention 2 [2606.03928].

At \(4\times\) KV-cache compression on Qwen3-4B, the average pass@1 accuracies are reported as follows: Full model \(65.0\), SeerAttention-R \(58.8\), SnapKV \(49.2\), R-KV \(54.7\), CurDKV \(49.8\), VaSE–AttnV \(57.5\), and VaSE–DKV \(59.1\). On Qwen3-14B, VaSE–DKV achieves \(65.8\%\), matching SeerAttention-R at \(65.4\%\) and exceeding R-KV at \(60.9\%\). The paper summarizes this by stating that, across six reasoning tasks, Qwen3 models using VaSE with \(4\times\) KV-cache compression yield higher average accuracies than the SOTA selection method at the same sparsity, while outperforming the strongest eviction method by more than \(4\%\) [2606.03928].

The systems results emphasize static-memory inference. Full caching scales as \(O(T)\) in keys and values and reaches OOM beyond 32 K tokens, whereas eviction methods have static overhead approximately \(K\cdot 2d\cdot{\tt sizeof(float)}\). Throughput scales inversely with \(K\); at \(K=2048\), VaSE–DKV processes 450 tok/s versus 140 tok/s for the full model on 14B. Reported peak memory is approximately 60 GB for the full model, compared with approximately 30 GB for VaSE–DKV and approximately 32 GB for VaSE–AttnV. The method is therefore positioned as a way to support FlashAttention2 while maintaining a static memory footprint for long-form reasoning [2606.03928].

## 6. Related formulations, terminological scope, and open directions

A source of terminological ambiguity is that “VaSE” appears in two distinct senses in the 2026 literature. In "Value-Aware Stochastic KV Cache Eviction for Reasoning Models" [2606.03928], VaSE denotes a specific recipe: reserve the top-\(v\) value-magnitude tokens and sample the rest. In "Forget Without Compromise: Nexus Sampling for Streaming KV-Cache Eviction Under Fixed Budgets" [2606.23961], “Value-aware stochastic eviction (VaSE)” is presented more broadly as a design prescription consisting of value estimation, stochastic retention, and budget compliance. Nexus Sampling is then described as an instantiation of that prescription through a two-term Nexus score \(s_j=a_j+\lambda\tilde c_j\), weighted reservoir sampling, and exact fixed-budget selection [2606.23961].

This broader framing helps situate the original VaSE recipe among neighboring approaches. VECTOR, for example, is a deterministic plug-in for eviction-based pipelines that performs three-way token routing—retention, approximation, and eviction—using both a base importance score and an offline-calibrated reconstructability signal for values. It retains all keys, reconstructs selected values on the fly with a layerwise OLS predictor, and is explicitly non-random once \(p_c\), \(p_a\), and \(W_{\mathrm{OLS}}\) are fixed [2605.23258]. Relative to VaSE, VECTOR addresses information loss through approximation rather than through stochastic retention and protected reservation.

The open problems identified for the original VaSE recipe are narrowly scoped and technically specific. The work focuses on decode-phase eviction, leaving extension to prefill-phase compression to future work; it is evaluated only on Qwen3, so application to other architectures remains to be validated; and it points to a possible interaction with quantization, noting that large-magnitude values also drive quantization error because the scaling factor \(s_i=\mathrm{Range}(v_i)/(2^b-1)\) becomes large, implying coarse reconstruction. The paper further identifies mechanistic interpretability of these outlier values—described as possible loop-breakers in chain-of-thought transitions—as an unresolved question [2606.03928].

Within this literature, VaSE is therefore best understood not merely as a cache-thinning heuristic, but as a precise response to two empirically observed weaknesses of deterministic eviction: the fragility induced by irreversible top-\(K\) cutoffs and the disproportionate importance of a small set of high-magnitude value states.

Source: https://www.emergentmind.com/topics/value-aware-stochastic-kv-cache-eviction-vase