Papers
Topics
Authors
Recent
Search
2000 character limit reached

KVSink: Enhancing KV Quantization in LLMs

Updated 17 December 2025
  • KVSink is a plug-and-play methodology for KV cache quantization that preserves critical attention sinks—token positions absorbing disproportionate attention mass.
  • The approach dynamically identifies sink tokens at a fixed emergence layer, excluding them from quantization to recover over 95% of FP16 perplexity.
  • Empirical evaluations on models like LLaMA2 show that KVSink outperforms PFN, offering robust performance and minimal computational overhead for aggressive model compression.

KVSink is a plug-and-play methodology for key-value (KV) cache quantization in LLM inference that targets the preservation of “attention sinks”—token positions that disproportionately absorb attention mass and introduce quantization sensitivity. It advances prior quantization pipelines by algorithmically predicting sink tokens during inference, permitting targeted exclusion from quantization and thereby maximizing perplexity (PPL) recovery with minimal overhead (Su et al., 6 Aug 2025).

1. Definition and Role of Attention Sinks

An attention sink is formally a token position ii that consistently aggregates an outsized fraction of attention mass in multiple heads and layers. For a single head at layer ll, attention scores are defined by

et,il,k=⟨qtl,k,kil,k⟩dk,pt,il,k=Softmaxi(et,il,k),e^{l,k}_{t,i} = \frac{\langle q^{l,k}_t, k^{l,k}_i \rangle}{\sqrt{d_k}},\quad p^{l,k}_{t,i} = \mathrm{Softmax}_i(e^{l,k}_{t,i}),

and the set of sinks for that head is

Sl,k={i ∣ ∀t,  pt,il,k≫1t}.S^{l,k} = \left\{ i~|~\forall t,\; p^{l,k}_{t,i} \gg \frac{1}{t} \right\}.

Aggregating across all heads and layers,

S=⋃l,kSl,k.S = \bigcup_{l,k} S^{l,k}.

Sink tokens i∈Si \in S frequently show numerically small vil,kv^{l,k}_i due to QKV suppression, so quantization error in VV-cache at these indices yields disproportionately large errors in attention outputs. Formally, the attention output can be decomposed as

Attention(Q,K,V)t=∑i∉Spitvi⏟ordinary+∑i∈Spitvi⏟sink bias,\mathrm{Attention}(Q,K,V)^t = \underbrace{\sum_{i \notin S} p_i^t v_i}_{\text{ordinary}} + \underbrace{\sum_{i \in S} p_i^t v_i}_{\text{sink bias}},

where the sink bias term is nearly constant across tt and highly sensitive to quantization of the corresponding ll0.

2. Cross-Layer Dynamics of Extreme Activation Outliers

The emergence and propagation of extreme channel-wise outliers underlie the formation of attention sinks. Tracking the tensors per block ll1: ll2

ll3

ll4

one observes five distinct stages in outlier magnitude and position:

  • Initial: Absence of extreme outliers.
  • Emergence (layer ll5): Spikes in ll6 propagate to ll7, creating stable outliers in ll8.
  • Stabilization: Intermediate layers sustain large, consistent outliers in ll9 and et,il,k=⟨qtl,k,kil,k⟩dk,pt,il,k=Softmaxi(et,il,k),e^{l,k}_{t,i} = \frac{\langle q^{l,k}_t, k^{l,k}_i \rangle}{\sqrt{d_k}},\quad p^{l,k}_{t,i} = \mathrm{Softmax}_i(e^{l,k}_{t,i}),0, while et,il,k=⟨qtl,k,kil,k⟩dk,pt,il,k=Softmaxi(et,il,k),e^{l,k}_{t,i} = \frac{\langle q^{l,k}_t, k^{l,k}_i \rangle}{\sqrt{d_k}},\quad p^{l,k}_{t,i} = \mathrm{Softmax}_i(e^{l,k}_{t,i}),1 and et,il,k=⟨qtl,k,kil,k⟩dk,pt,il,k=Softmaxi(et,il,k),e^{l,k}_{t,i} = \frac{\langle q^{l,k}_t, k^{l,k}_i \rangle}{\sqrt{d_k}},\quad p^{l,k}_{t,i} = \mathrm{Softmax}_i(e^{l,k}_{t,i}),2 subside.
  • Dissipation: Counter-spikes in et,il,k=⟨qtl,k,kil,k⟩dk,pt,il,k=Softmaxi(et,il,k),e^{l,k}_{t,i} = \frac{\langle q^{l,k}_t, k^{l,k}_i \rangle}{\sqrt{d_k}},\quad p^{l,k}_{t,i} = \mathrm{Softmax}_i(e^{l,k}_{t,i}),3 nullify prior stable outliers.
  • Final: All activations return to baseline magnitudes.

These persistent, stable outliers correspond to emerging and sustained attention sinks, which significantly influence downstream attention and the sensitivity to quantization.

3. Limitations of the Preserve-First-N Strategy

The Preserve-First-N (PFN) strategy excludes the first et,il,k=⟨qtl,k,kil,k⟩dk,pt,il,k=Softmaxi(et,il,k),e^{l,k}_{t,i} = \frac{\langle q^{l,k}_t, k^{l,k}_i \rangle}{\sqrt{d_k}},\quad p^{l,k}_{t,i} = \mathrm{Softmax}_i(e^{l,k}_{t,i}),4 tokens’ KVs from quantization. There are two core limitations:

  • Attention sinks may occur at positions beyond et,il,k=⟨qtl,k,kil,k⟩dk,pt,il,k=Softmaxi(et,il,k),e^{l,k}_{t,i} = \frac{\langle q^{l,k}_t, k^{l,k}_i \rangle}{\sqrt{d_k}},\quad p^{l,k}_{t,i} = \mathrm{Softmax}_i(e^{l,k}_{t,i}),5 (e.g., token 14 in LLaMA2-7B prefill). PFN cannot detect or exclude such outliers.
  • Empirical evidence demonstrates sudden PPL degradation if a non-excluded sink emerges after et,il,k=⟨qtl,k,kil,k⟩dk,pt,il,k=Softmaxi(et,il,k),e^{l,k}_{t,i} = \frac{\langle q^{l,k}_t, k^{l,k}_i \rangle}{\sqrt{d_k}},\quad p^{l,k}_{t,i} = \mathrm{Softmax}_i(e^{l,k}_{t,i}),6, regardless of how many initial tokens are preserved. Thus, PFN is not robust against atypical sink locations and cannot guarantee PPL stability (Su et al., 6 Aug 2025).

4. KVSink: Algorithmic Sink Token Identification and Preservation

KVSink mitigates these limitations by dynamically predicting sink positions in the prefill phase through activation outlier identification at a fixed emergence layer et,il,k=⟨qtl,k,kil,k⟩dk,pt,il,k=Softmaxi(et,il,k),e^{l,k}_{t,i} = \frac{\langle q^{l,k}_t, k^{l,k}_i \rangle}{\sqrt{d_k}},\quad p^{l,k}_{t,i} = \mathrm{Softmax}_i(e^{l,k}_{t,i}),7. The algorithm:

  • Selects a pre-identified outlier channel et,il,k=⟨qtl,k,kil,k⟩dk,pt,il,k=Softmaxi(et,il,k),e^{l,k}_{t,i} = \frac{\langle q^{l,k}_t, k^{l,k}_i \rangle}{\sqrt{d_k}},\quad p^{l,k}_{t,i} = \mathrm{Softmax}_i(e^{l,k}_{t,i}),8 at layer et,il,k=⟨qtl,k,kil,k⟩dk,pt,il,k=Softmaxi(et,il,k),e^{l,k}_{t,i} = \frac{\langle q^{l,k}_t, k^{l,k}_i \rangle}{\sqrt{d_k}},\quad p^{l,k}_{t,i} = \mathrm{Softmax}_i(e^{l,k}_{t,i}),9.
  • For the hidden-state tensor Sl,k={i ∣ ∀t,  pt,il,k≫1t}.S^{l,k} = \left\{ i~|~\forall t,\; p^{l,k}_{t,i} \gg \frac{1}{t} \right\}.0, determines the threshold Sl,k={i ∣ ∀t,  pt,il,k≫1t}.S^{l,k} = \left\{ i~|~\forall t,\; p^{l,k}_{t,i} \gg \frac{1}{t} \right\}.1 as the Sl,k={i ∣ ∀t,  pt,il,k≫1t}.S^{l,k} = \left\{ i~|~\forall t,\; p^{l,k}_{t,i} \gg \frac{1}{t} \right\}.2-th largest Sl,k={i ∣ ∀t,  pt,il,k≫1t}.S^{l,k} = \left\{ i~|~\forall t,\; p^{l,k}_{t,i} \gg \frac{1}{t} \right\}.3 over token positions.
  • The predicted sink set Sl,k={i ∣ ∀t,  pt,il,k≫1t}.S^{l,k} = \left\{ i~|~\forall t,\; p^{l,k}_{t,i} \gg \frac{1}{t} \right\}.4 comprises all positions Sl,k={i ∣ ∀t,  pt,il,k≫1t}.S^{l,k} = \left\{ i~|~\forall t,\; p^{l,k}_{t,i} \gg \frac{1}{t} \right\}.5 where Sl,k={i ∣ ∀t,  pt,il,k≫1t}.S^{l,k} = \left\{ i~|~\forall t,\; p^{l,k}_{t,i} \gg \frac{1}{t} \right\}.6.

Pseudocode: i∈Si \in S4

The quantization step uses: Sl,k={i ∣ ∀t,  pt,il,k≫1t}.S^{l,k} = \left\{ i~|~\forall t,\; p^{l,k}_{t,i} \gg \frac{1}{t} \right\}.7

Computational complexity is dominated by a single top-Sl,k={i ∣ ∀t,  pt,il,k≫1t}.S^{l,k} = \left\{ i~|~\forall t,\; p^{l,k}_{t,i} \gg \frac{1}{t} \right\}.8 operation (Sl,k={i ∣ ∀t,  pt,il,k≫1t}.S^{l,k} = \left\{ i~|~\forall t,\; p^{l,k}_{t,i} \gg \frac{1}{t} \right\}.9), negligible for contemporary context lengths (S=⋃l,kSl,k.S = \bigcup_{l,k} S^{l,k}.0, S=⋃l,kSl,k.S = \bigcup_{l,k} S^{l,k}.1).

5. Experimental Validation and Comparative Performance

Experiments span models including LLaMA2-7B/13B/70B, LLaMA2-7B-chat, Mistral-7B, LLaMA3-8B, and LLaMA3.1-8B-instruct, evaluated on Wikitext-2 and C4 with variable quantization schemes (RTN INT2/INT4, static/dynamic per-token and per-channel). Baselines are PFN(S=⋃l,kSl,k.S = \bigcup_{l,k} S^{l,k}.2) and KVQuant.

Results Overview

Model/Method FP16 PPL PFN(5) PPL KVSink(5) PPL
LLaMA2-70B, 4-bit 2.5 59.5 5.0

Across all tested LLMs, preserving S=⋃l,kSl,k.S = \bigcup_{l,k} S^{l,k}.3 sinks via KVSink recovers S=⋃l,kSl,k.S = \bigcup_{l,k} S^{l,k}.4 of FP16 PPL, while PFN requires S=⋃l,kSl,k.S = \bigcup_{l,k} S^{l,k}.5 and still fails in some cases.

For KVSink + KVQuant integration (Wikitext-2, LLaMA2-7B, 2-bit quantization, 1% outlier isolation):

  • KVQuant PPL: 5.53
  • KVSink-5 PPL: 5.44

On C4 dataset: 6.94 S=⋃l,kSl,k.S = \bigcup_{l,k} S^{l,k}.6 6.81 (KVQuant S=⋃l,kSl,k.S = \bigcup_{l,k} S^{l,k}.7 KVSink).

Importantly, lowering the outlier-isolation budget (to 0.1%) does not hurt PPL with KVSink, enabling more aggressive model compression.

6. Guidelines for Practical Integration

KVSink integration follows static configuration per architecture:

  • Identify S=⋃l,kSl,k.S = \bigcup_{l,k} S^{l,k}.8 once for each model (cf. Table 9).
  • Target S=⋃l,kSl,k.S = \bigcup_{l,k} S^{l,k}.9 sink tokens (empirically sufficient for tested LLMs).
  • In static-quant pipelines, also exclude sink set i∈Si \in S0 from calibration.
  • Implementation overhead is minimal: i∈Si \in S10.04–0.05 ms latency per 4K-token sequence (A100), i∈Si \in S2 MB memory for storage.
  • Recommended practice: wrap quantizer, perform outlier detection at i∈Si \in S3, and “freeze” sink set for remainder of inference.

KVSink thus provides a methodologically grounded, low-cost approach for attention sink preservation, consistently improving quantized inference accuracy over both PFN and state-of-the-art alternatives (Su et al., 6 Aug 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to KVSink.