---
title: 'HySparse2 Long-Context Architectured: Hybrid with KV'
url: https://www.emergentmind.com/topics/hysparse2
type: topic
---

# HySparse2 Long-Context Architectured: Hybrid with KV

HySparse2 is a long-context Transformer architecture for agentic inference that combines hybrid sparse attention with two-level key–value (KV) sharing [2609.26368]. Its outer mechanism, **KV Bridging**, adopts a YOCO-style self-decoder/cross-decoder organization and constructs cross-decoder full-attention KV caches from self-decoder hidden states. Its inner mechanism, **KV Reuse**, shares full-attention KV caches and token-selection indices with subsequent sparse-attention layers. HySparse2 further replaces block-level sparsity with token-level selection and incorporates a forced recent-token window directly into sparse attention. The resulting architecture enables prefill to terminate after the self-decoder while preserving cross-decoder execution during autoregressive decoding.

## 1. Workload and architectural motivation

HySparse2 targets long-horizon, multi-turn agentic workloads in which short actions or tool calls are followed by long observations, such as search results, documents, tool outputs, execution traces, and intermediate reasoning. Each new observation expands the accumulated context, creating simultaneous demands for efficient prefill, compact KV-cache storage, and accurate long-context retrieval.

Conventional full attention has quadratic prefill cost with respect to sequence length. For queries $Q$, keys $K$, values $V$, and causal mask $M$, attention is represented schematically as

$$
A(Q,K,V)=\operatorname{softmax}\left(\frac{QK^\top}{\sqrt d}+M\right)V.
$$

The prefill score matrix contains $O(L^2)$ entries for sequence length $L$. During decoding, attention to a cached history remains approximately $O(Ld)$ per generated token, while storing K/V states for all layers grows linearly with $L$ but includes the multiplicative costs of layers, KV heads, head dimensions, and numerical precision.

HySparse2 does not remove global attention entirely. Instead, a small number of full-attention layers perform global mixing and provide attention scores for token selection. Subsequent sparse layers reuse the resulting KV representations and selected indices. This design allocates exact global computation to index-generating layers while limiting later attention operations to selected tokens.

The architecture is motivated particularly by contexts in which relevant evidence is distributed across multiple turns and may occur at arbitrary token positions. Token-level selection avoids retaining entire fixed-size blocks when only one or a few tokens in those blocks are relevant.

## 2. Two-level KV sharing

HySparse2 has two complementary levels of KV sharing.

**Outer level—KV Bridging:** the model is divided into a self-decoder and a cross-decoder. Full-attention layers in the self-decoder provide the hidden states from which the K/V caches of full-attention layers in the cross-decoder are generated.

**Inner level—KV Reuse:** within hybrid blocks, a full-attention layer constructs a full KV cache and token-selection indices. Subsequent sparse-attention layers reuse the selected KV entries and the corresponding indices instead of constructing independent full-length caches.

The combined structure is:

$$
\text{self-decoder hidden states}
\xrightarrow{\text{KV Bridging}}
\text{cross-decoder full-attention KV caches}
$$

and

$$
\text{full-attention KV cache and index}
\xrightarrow{\text{KV Reuse}}
\text{sparse-layer KV and support}.
$$

These mechanisms address different forms of redundancy. KV Bridging eliminates the need to generate cross-decoder KV caches from cross-decoder prefill hidden states. KV Reuse prevents multiple sparse layers following the same full-attention layer from independently materializing equivalent or overlapping long-context representations.

## 3. KV Bridging and prefill early exit

The 49-layer HySparse2 model is divided into a self-decoder and a cross-decoder. The self-decoder uses full attention and sliding-window attention (SWA), whereas the cross-decoder uses full attention and sparse attention.

For a self-decoder full-attention layer $i$ and a cross-decoder full-attention layer $j$, the bridged caches are

$$
\mathbf K^{\mathrm{cross}_j}
=
\operatorname{Proj}^{K}_{j}
\left(\mathbf H^{\mathrm{self}_i}\right),
$$

$$
\mathbf V^{\mathrm{cross}_j}
=
\operatorname{Proj}^{V}_{j}
\left(\mathbf H^{\mathrm{self}_i}\right),
$$

while the cross-decoder query is computed from its own hidden state:

$$
\mathbf Q^{\mathrm{cross}_j}
=
\operatorname{Proj}^{Q}_{j}
\left(\mathbf H^{\mathrm{cross}_j}\right).
$$

Thus, KV Bridging shares the source representation for keys and values but does not share queries. Each cross-decoder full-attention layer has its own K/V projections, even when several layers use hidden states originating from the same self-decoder layer.

Only full-attention layers are bridged. The paper’s stated rationale is that full-attention hidden states incorporate global information and are therefore more suitable for generating cross-decoder KV caches than SWA states. One self-decoder full-attention layer may supply multiple cross-decoder full-attention layers, allowing the self- and cross-decoders to contain different proportions of full-attention layers.

During prefill, the execution is:

1. The self-decoder processes the entire newly prefilling sequence.
2. Self-decoder full-attention states are projected into cross-decoder full-attention K/V caches.
3. Full-attention layers provide token-selection information for sparse layers.
4. Sparse layers reuse the bridged caches and selection metadata.
5. Prefill terminates without executing the cross-decoder layers over the entire sequence.

The cross-decoder remains active during autoregressive decoding. It is skipped only during prefill cache construction. In the evaluated 49-layer configuration, prefill requires the first 25 layers—the self-decoder and bridging projections—instead of all 49 layers.

A separate SWA branch in the cross-decoder would undermine this early exit because its K/V states would depend on cross-decoder hidden states, which themselves depend recursively on preceding cross-decoder layers. HySparse2 removes that dependency by eliminating the separate cross-decoder SWA branch.

## 4. Hybrid attention and token selection

The self-decoder combines full attention with SWA. For position $t$ and window width $w$, the local support is

$$
\mathcal W_t=
\{\max(1,t-w+1),\ldots,t\}.
$$

SWA restricts attention to this set:

$$
\operatorname{SWA}(q_t,K,V)
=
\operatorname{softmax}
\left(
\frac{q_tK_{\mathcal W_t}^{\top}}{\sqrt d}
\right)
V_{\mathcal W_t}.
$$

The evaluated configuration uses a 128-token window. Self-decoder SWA layers use partial RoPE with 64 rotary dimensions and base $10{,}000$. Full-attention and sparse-attention layers use NoPE.

The cross-decoder contains full-attention indexer layers and sparse-attention layers. For query $q_t$, sparse attention operates on a selected token set $\mathcal S_t$:

$$
\operatorname{SA}(q_t,K,V)
=
\operatorname{softmax}
\left(
\frac{q_tK_{\mathcal S_t}^{\top}}{\sqrt d}
\right)
V_{\mathcal S_t}.
$$

The selection scores are the full-attention logits,

$$
s_{t,\ell}=\frac{q_t^\top k_\ell}{\sqrt d},
$$

subject to the causal constraint $\ell\leq t$. Subsequent sparse layers reuse the full-attention layer’s selected indices and KV entries.

### Token-level sparsity

HySparse selected 1,024 global tokens in 64-token blocks. HySparse2 selects individual tokens. In the evaluated configuration, 128 recent local tokens are forced into the support and 1,024 additional global tokens are selected by score:

$$
\mathcal W_t^{\mathrm{recent}}
=
\{\max(1,t-127),\ldots,t\},
$$

$$
\mathcal G_t
=
\operatorname{TopK}_{\ell\notin\mathcal W_t^{\mathrm{recent}}}
(s_{t,\ell},1024),
$$

$$
\mathcal S_t
=
\mathcal W_t^{\mathrm{recent}}\cup\mathcal G_t.
$$

Token-level selection permits the budget to be distributed across multiple turns without retaining up to 63 neighboring tokens for each selected token, as would occur with 64-token blocks.

The token-versus-block ablation reported the following results:

| Metric | Block selection | Token selection |
|---|---:|---:|
| RULER-v2 | 49.56 | 56.13 |
| MRCR-v2, two needles | 12.94 | 21.08 |
| GraphWalks | 29.38 | 34.92 |
| MMLU-Pro | 35.74 | 36.97 |
| NoLiMa | 40.27 | 38.43 |
| BBH | 61.93 | 60.70 |

Token-level sparsity improved the principal long-context retrieval metrics but was not uniformly superior on every task.

### Forced recent-token selection

HySparse2 removes the separate SWA branch from sparse layers. Instead, recent tokens are inserted directly into the same sparse support:

$$
\mathcal S_t
=
\mathcal W_t^{\mathrm{recent}}
\cup
\operatorname{TopK}
\left(
s_{t,\ell}:
\ell\notin\mathcal W_t^{\mathrm{recent}}
\right).
$$

Local and globally selected tokens are therefore read from the same full-attention KV cache. There is no separate local K/V projection, second local cache, or gated fusion branch.

The local-attention ablation compared Gated SWA, No SWA, and Forced SWA:

| Metric | Gated SWA | No SWA | Forced SWA |
|---|---:|---:|---:|
| RULER | 88.19 | 84.55 | **89.84** |
| RULER-v2 | 53.66 | 54.62 | **55.98** |
| MRCR-v2 | **27.66** | 20.73 | 22.67 |
| GraphWalks | 35.39 | 36.48 | **37.13** |
| LongPPL | 6.8807 | 7.1307 | 6.9838 |
| GSM8K | **64.52** | 60.35 | 59.44 |

Forced SWA enabled the early-exit prefill structure and performed strongly on several retrieval metrics, but it underperformed Gated SWA on MRCR-v2 and GSM8K and had higher LongPPL.

## 5. Model configuration and inference workflow

The evaluated models are 80B-A3B mixture-of-experts models with:

- 49 Transformer layers;
- hidden size 2,048;
- approximately 80B total parameters and 3B active parameters;
- FP8 KV-cache storage;
- simplified mHC residual mixing with the residual matrix fixed to identity.

The attention configurations are:

| Model | Full-attention layers | Q/KV heads | QK/V head dimension |
|---|---:|---:|---:|
| Hybrid SWA | 9 | 64/4 | 192/128 |
| HySparse | 5 | 64/4 | 192/128 |
| HySparse2 | 5 | 64/1 | 256/256 |

HySparse2 uses multi-query attention (MQA), whereas the baselines use grouped-query attention (GQA). This difference contributes to the KV-cache reduction and affects the comparison independently of the attention sparsity pattern.

The principal HySparse2 settings are:

- 128 forced local tokens;
- 1,024 global tokens;
- token-level selection;
- self-decoder SWA;
- cross-decoder sparse attention;
- NoPE in full and sparse attention;
- partial RoPE in self-decoder SWA;
- sigmoid output gates in sparse/SWA layers;
- learnable per-head sink biases.

The training setup uses approximately 500B pretraining tokens with 32k pretraining context and approximately 100B post-training tokens. Agentic data are added during post-training, and the context is extended to 256k. The compared models use the same data and schedules, the Muon optimizer, and a WSD schedule. Peak learning rates are $10^{-3}$ for pretraining and $5\times10^{-5}$ for post-training.

For a prefilling sequence, the workflow is:

1. Run the self-decoder using full attention and SWA.
2. Extract hidden states from self-decoder full-attention layers.
3. Project those states into cross-decoder full-attention K/V caches.
4. Generate full-attention token-selection metadata.
5. Add the forced recent-token window to each sparse support.
6. Transfer the bridged caches and selection metadata to decoding.
7. End prefill after the self-decoder and bridging operations.
8. Run the cross-decoder normally during autoregressive decoding.

During decoding, cross-decoder queries are computed from cross-decoder hidden states. Full-attention layers use bridged K/V caches, while sparse layers access selected entries from those caches.

## 6. Evaluation, efficiency, and limitations

### Quality results

Selected pretraining results are:

| Task | Hybrid SWA | HySparse | HySparse2 |
|---|---:|---:|---:|
| RULER | 88.71 | 84.89 | **90.77** |
| NoLiMa | 30.13 | 40.27 | **49.76** |
| Repo Code PPL | 1.1578 | 1.1588 | **1.1570** |
| BBH | 60.14 | 61.93 | **64.29** |
| MMLU-Pro | 36.12 | 35.74 | **37.56** |
| DROP | 60.90 | **63.78** | 58.99 |
| GSM8K | 64.44 | 64.14 | 61.94 |

HySparse2 shows its strongest advantage on long-context evaluation but underperforms HySparse on DROP and GSM8K.

The post-training evaluation includes AgentPPL, LongPPL, MRCR-v2, RULER-v2, and GraphWalks. Relative to HySparse, HySparse2 improves mean MRCR-v2 by 11.30 percentage points and mean RULER-v2 by 19.81 points. Relative to Hybrid SWA, the corresponding improvements are 6.44 and 18.65 points.

At 256k context, RULER-v2 scores are:

- HySparse2: 58.45;
- HySparse: 32.61;
- Hybrid SWA: 35.74.

The paper reports lower AgentPPL and LongPPL for HySparse2 at all evaluated lengths. AgentPPL increases with context length, whereas LongPPL decreases, reflecting the different effects of longer context on multi-turn retrieval and long-range-dependent tokens.

### Prefill computation and KV storage

At one million tokens with FP8 KV storage, the reported cache sizes are:

| Quantity | Hybrid SWA | HySparse | HySparse2 |
|---|---:|---:|---:|
| KV cache | 12.09 GB | 6.72 GB | **2.69 GB** |

Reported prefill FLOP reductions at one million tokens are 2.92 times relative to HySparse and 5.02 times relative to Hybrid SWA. The reduction results from:

- terminating prefill after the self-decoder;
- using only one self-decoder full-attention layer during prefill in the evaluated configuration;
- sharing cross-decoder KV caches;
- reusing full-attention KV entries in sparse layers;
- removing the separate sparse-layer SWA branch;
- using MQA instead of GQA;
- selecting only a limited token support in sparse attention.

A conventional KV-storage estimate is proportional to

$$
M_{\mathrm{KV}}
\propto
L\sum_{\ell}
n_{\mathrm{KV},\ell}
d_{\mathrm{head},\ell}b,
$$

where $L$ is sequence length, $n_{\mathrm{KV},\ell}$ is the number of KV heads, $d_{\mathrm{head},\ell}$ is the head dimension, and $b$ is bytes per element. HySparse2 reduces this quantity through MQA, cross-layer sharing, and sparse reuse.

### KV Bridging ablations

A separate KV Bridging study at 290B-A8B scale used approximately 1.8T training tokens and 32k context:

| Task | Without bridging | With bridging |
|---|---:|---:|
| MMLU | 72.68 | 72.80 |
| TriviaQA | 73.32 | 74.10 |
| BBH | 70.65 | 69.57 |
| DROP | 71.37 | 68.17 |
| GSM8K | 77.63 | 76.65 |
| Repo Code PPL | 1.1351 | 1.1353 |
| RULER | 96.32 | 96.01 |
| LongPPL | 3.6053 | 3.4202 |

KV Bridging produces broadly comparable quality, with improvements on MMLU, TriviaQA, and LongPPL but declines on BBH, DROP, and GSM8K. In a separate comparison with KV Mirror, the final RULER scores are 87.65 for KV Bridging and 81.28 for KV Mirror.

### Limitations

HySparse2 retains several costs and risks:

- **Selection errors:** relevant tokens not selected by full-attention scores are inaccessible to sparse layers.
- **Residual full-attention cost:** five full-attention layers remain in the full model, and one self-decoder full-attention layer is still used during prefill.
- **Irregular token access:** token-level sparsity can be less hardware-friendly than regular block sparsity.
- **Budget competition:** forced recent-token inclusion consumes part of the sparse support available for globally retrieved tokens.
- **Task-dependent local modeling:** Forced SWA is not uniformly superior to a separate gated SWA branch.
- **Representational restriction:** bridged K/V states are projections of self-decoder full-attention states rather than cross-decoder-native hidden states.
- **Confounded comparisons:** HySparse2 uses MQA while HySparse and Hybrid SWA use GQA, so cache and efficiency differences do not arise solely from the sparsity pattern.
- **Mixed general reasoning:** long-context retrieval improves substantially, but some general reasoning metrics decline.
- **Incomplete end-to-end throughput evidence:** the supplied results report cache sizes and prefill FLOP reductions, but not a comprehensive decoding latency or tokens-per-second table.

HySparse2 is therefore best characterized as a coordinated architecture for long-context agentic inference rather than merely a sparse-attention kernel. Its defining pipeline is

$$
\text{self-decoder full/SWA processing}
\rightarrow
\text{KV Bridging}
\rightarrow
\text{cross-decoder full-attention indexing}
\rightarrow
\text{forced-window token selection}
\rightarrow
\text{KV Reuse in sparse layers}.
$$

The central systems consequence is that prefill can terminate after the self-decoder, while the central modeling consequence is that sparse layers retain both recent local context and globally selected individual tokens.

Source: https://www.emergentmind.com/topics/hysparse2