HySparse2 Long-Context Architectured: Hybrid with KV
- HySparse2 is a long-context transformer architecture designed for agentic inference, which leverages hybrid sparse attention with KV Bridging and KV Reuse mechanisms that optimize full-attention layers, and cross-decoder KV caches, ultimately improving efficiency by reducing cache storage and computational iteration in tasks requiring length and contextual integrity,.
- HySparse2 addresses long-horizon, multi-turn agentic workloads by efficiently managing context through KV Bridging and KV Reuse. KV Bridging minimizes redundancy between the self- and cross-decoder. KV Reuse allows sparse layers to efficiently reuse key-value caches and token selections.
- Token-level selection in HySparse2 improves sparse attention with a 128 recent local token forced window and 1,024 global selected tokens, aiding performance in benchmarks such as RULER-v2 and MRCR-v2.
HySparse2 is a long-context Transformer architecture for agentic inference that combines hybrid sparse attention with two-level key–value (KV) sharing (Wei et al., 22 Sep 2026). Its outer mechanism, KV Bridging, adopts a YOCO-style self-decoder/cross-decoder organization and constructs cross-decoder full-attention KV caches from self-decoder hidden states. Its inner mechanism, KV Reuse, shares full-attention KV caches and token-selection indices with subsequent sparse-attention layers. HySparse2 further replaces block-level sparsity with token-level selection and incorporates a forced recent-token window directly into sparse attention. The resulting architecture enables prefill to terminate after the self-decoder while preserving cross-decoder execution during autoregressive decoding.
1. Workload and architectural motivation
HySparse2 targets long-horizon, multi-turn agentic workloads in which short actions or tool calls are followed by long observations, such as search results, documents, tool outputs, execution traces, and intermediate reasoning. Each new observation expands the accumulated context, creating simultaneous demands for efficient prefill, compact KV-cache storage, and accurate long-context retrieval.
Conventional full attention has quadratic prefill cost with respect to sequence length. For queries , keys , values , and causal mask , attention is represented schematically as
The prefill score matrix contains entries for sequence length . During decoding, attention to a cached history remains approximately per generated token, while storing K/V states for all layers grows linearly with but includes the multiplicative costs of layers, KV heads, head dimensions, and numerical precision.
HySparse2 does not remove global attention entirely. Instead, a small number of full-attention layers perform global mixing and provide attention scores for token selection. Subsequent sparse layers reuse the resulting KV representations and selected indices. This design allocates exact global computation to index-generating layers while limiting later attention operations to selected tokens.
The architecture is motivated particularly by contexts in which relevant evidence is distributed across multiple turns and may occur at arbitrary token positions. Token-level selection avoids retaining entire fixed-size blocks when only one or a few tokens in those blocks are relevant.
2. Two-level KV sharing
HySparse2 has two complementary levels of KV sharing.
Outer level—KV Bridging: the model is divided into a self-decoder and a cross-decoder. Full-attention layers in the self-decoder provide the hidden states from which the K/V caches of full-attention layers in the cross-decoder are generated.
Inner level—KV Reuse: within hybrid blocks, a full-attention layer constructs a full KV cache and token-selection indices. Subsequent sparse-attention layers reuse the selected KV entries and the corresponding indices instead of constructing independent full-length caches.
The combined structure is:
and
0
These mechanisms address different forms of redundancy. KV Bridging eliminates the need to generate cross-decoder KV caches from cross-decoder prefill hidden states. KV Reuse prevents multiple sparse layers following the same full-attention layer from independently materializing equivalent or overlapping long-context representations.
3. KV Bridging and prefill early exit
The 49-layer HySparse2 model is divided into a self-decoder and a cross-decoder. The self-decoder uses full attention and sliding-window attention (SWA), whereas the cross-decoder uses full attention and sparse attention.
For a self-decoder full-attention layer 1 and a cross-decoder full-attention layer 2, the bridged caches are
3
4
while the cross-decoder query is computed from its own hidden state:
5
Thus, KV Bridging shares the source representation for keys and values but does not share queries. Each cross-decoder full-attention layer has its own K/V projections, even when several layers use hidden states originating from the same self-decoder layer.
Only full-attention layers are bridged. The paper’s stated rationale is that full-attention hidden states incorporate global information and are therefore more suitable for generating cross-decoder KV caches than SWA states. One self-decoder full-attention layer may supply multiple cross-decoder full-attention layers, allowing the self- and cross-decoders to contain different proportions of full-attention layers.
During prefill, the execution is:
- The self-decoder processes the entire newly prefilling sequence.
- Self-decoder full-attention states are projected into cross-decoder full-attention K/V caches.
- Full-attention layers provide token-selection information for sparse layers.
- Sparse layers reuse the bridged caches and selection metadata.
- Prefill terminates without executing the cross-decoder layers over the entire sequence.
The cross-decoder remains active during autoregressive decoding. It is skipped only during prefill cache construction. In the evaluated 49-layer configuration, prefill requires the first 25 layers—the self-decoder and bridging projections—instead of all 49 layers.
A separate SWA branch in the cross-decoder would undermine this early exit because its K/V states would depend on cross-decoder hidden states, which themselves depend recursively on preceding cross-decoder layers. HySparse2 removes that dependency by eliminating the separate cross-decoder SWA branch.
4. Hybrid attention and token selection
The self-decoder combines full attention with SWA. For position 6 and window width 7, the local support is
8
SWA restricts attention to this set:
9
The evaluated configuration uses a 128-token window. Self-decoder SWA layers use partial RoPE with 64 rotary dimensions and base 0. Full-attention and sparse-attention layers use NoPE.
The cross-decoder contains full-attention indexer layers and sparse-attention layers. For query 1, sparse attention operates on a selected token set 2:
3
The selection scores are the full-attention logits,
4
subject to the causal constraint 5. Subsequent sparse layers reuse the full-attention layer’s selected indices and KV entries.
Token-level sparsity
HySparse selected 1,024 global tokens in 64-token blocks. HySparse2 selects individual tokens. In the evaluated configuration, 128 recent local tokens are forced into the support and 1,024 additional global tokens are selected by score:
6
7
8
Token-level selection permits the budget to be distributed across multiple turns without retaining up to 63 neighboring tokens for each selected token, as would occur with 64-token blocks.
The token-versus-block ablation reported the following results:
| Metric | Block selection | Token selection |
|---|---|---|
| RULER-v2 | 49.56 | 56.13 |
| MRCR-v2, two needles | 12.94 | 21.08 |
| GraphWalks | 29.38 | 34.92 |
| MMLU-Pro | 35.74 | 36.97 |
| NoLiMa | 40.27 | 38.43 |
| BBH | 61.93 | 60.70 |
Token-level sparsity improved the principal long-context retrieval metrics but was not uniformly superior on every task.
Forced recent-token selection
HySparse2 removes the separate SWA branch from sparse layers. Instead, recent tokens are inserted directly into the same sparse support:
9
Local and globally selected tokens are therefore read from the same full-attention KV cache. There is no separate local K/V projection, second local cache, or gated fusion branch.
The local-attention ablation compared Gated SWA, No SWA, and Forced SWA:
| Metric | Gated SWA | No SWA | Forced SWA |
|---|---|---|---|
| RULER | 88.19 | 84.55 | 89.84 |
| RULER-v2 | 53.66 | 54.62 | 55.98 |
| MRCR-v2 | 27.66 | 20.73 | 22.67 |
| GraphWalks | 35.39 | 36.48 | 37.13 |
| LongPPL | 6.8807 | 7.1307 | 6.9838 |
| GSM8K | 64.52 | 60.35 | 59.44 |
Forced SWA enabled the early-exit prefill structure and performed strongly on several retrieval metrics, but it underperformed Gated SWA on MRCR-v2 and GSM8K and had higher LongPPL.
5. Model configuration and inference workflow
The evaluated models are 80B-A3B mixture-of-experts models with:
- 49 Transformer layers;
- hidden size 2,048;
- approximately 80B total parameters and 3B active parameters;
- FP8 KV-cache storage;
- simplified mHC residual mixing with the residual matrix fixed to identity.
The attention configurations are:
| Model | Full-attention layers | Q/KV heads | QK/V head dimension |
|---|---|---|---|
| Hybrid SWA | 9 | 64/4 | 192/128 |
| HySparse | 5 | 64/4 | 192/128 |
| HySparse2 | 5 | 64/1 | 256/256 |
HySparse2 uses multi-query attention (MQA), whereas the baselines use grouped-query attention (GQA). This difference contributes to the KV-cache reduction and affects the comparison independently of the attention sparsity pattern.
The principal HySparse2 settings are:
- 128 forced local tokens;
- 1,024 global tokens;
- token-level selection;
- self-decoder SWA;
- cross-decoder sparse attention;
- NoPE in full and sparse attention;
- partial RoPE in self-decoder SWA;
- sigmoid output gates in sparse/SWA layers;
- learnable per-head sink biases.
The training setup uses approximately 500B pretraining tokens with 32k pretraining context and approximately 100B post-training tokens. Agentic data are added during post-training, and the context is extended to 256k. The compared models use the same data and schedules, the Muon optimizer, and a WSD schedule. Peak learning rates are 0 for pretraining and 1 for post-training.
For a prefilling sequence, the workflow is:
- Run the self-decoder using full attention and SWA.
- Extract hidden states from self-decoder full-attention layers.
- Project those states into cross-decoder full-attention K/V caches.
- Generate full-attention token-selection metadata.
- Add the forced recent-token window to each sparse support.
- Transfer the bridged caches and selection metadata to decoding.
- End prefill after the self-decoder and bridging operations.
- Run the cross-decoder normally during autoregressive decoding.
During decoding, cross-decoder queries are computed from cross-decoder hidden states. Full-attention layers use bridged K/V caches, while sparse layers access selected entries from those caches.
6. Evaluation, efficiency, and limitations
Quality results
Selected pretraining results are:
| Task | Hybrid SWA | HySparse | HySparse2 |
|---|---|---|---|
| RULER | 88.71 | 84.89 | 90.77 |
| NoLiMa | 30.13 | 40.27 | 49.76 |
| Repo Code PPL | 1.1578 | 1.1588 | 1.1570 |
| BBH | 60.14 | 61.93 | 64.29 |
| MMLU-Pro | 36.12 | 35.74 | 37.56 |
| DROP | 60.90 | 63.78 | 58.99 |
| GSM8K | 64.44 | 64.14 | 61.94 |
HySparse2 shows its strongest advantage on long-context evaluation but underperforms HySparse on DROP and GSM8K.
The post-training evaluation includes AgentPPL, LongPPL, MRCR-v2, RULER-v2, and GraphWalks. Relative to HySparse, HySparse2 improves mean MRCR-v2 by 11.30 percentage points and mean RULER-v2 by 19.81 points. Relative to Hybrid SWA, the corresponding improvements are 6.44 and 18.65 points.
At 256k context, RULER-v2 scores are:
- HySparse2: 58.45;
- HySparse: 32.61;
- Hybrid SWA: 35.74.
The paper reports lower AgentPPL and LongPPL for HySparse2 at all evaluated lengths. AgentPPL increases with context length, whereas LongPPL decreases, reflecting the different effects of longer context on multi-turn retrieval and long-range-dependent tokens.
Prefill computation and KV storage
At one million tokens with FP8 KV storage, the reported cache sizes are:
| Quantity | Hybrid SWA | HySparse | HySparse2 |
|---|---|---|---|
| KV cache | 12.09 GB | 6.72 GB | 2.69 GB |
Reported prefill FLOP reductions at one million tokens are 2.92 times relative to HySparse and 5.02 times relative to Hybrid SWA. The reduction results from:
- terminating prefill after the self-decoder;
- using only one self-decoder full-attention layer during prefill in the evaluated configuration;
- sharing cross-decoder KV caches;
- reusing full-attention KV entries in sparse layers;
- removing the separate sparse-layer SWA branch;
- using MQA instead of GQA;
- selecting only a limited token support in sparse attention.
A conventional KV-storage estimate is proportional to
2
where 3 is sequence length, 4 is the number of KV heads, 5 is the head dimension, and 6 is bytes per element. HySparse2 reduces this quantity through MQA, cross-layer sharing, and sparse reuse.
KV Bridging ablations
A separate KV Bridging study at 290B-A8B scale used approximately 1.8T training tokens and 32k context:
| Task | Without bridging | With bridging |
|---|---|---|
| MMLU | 72.68 | 72.80 |
| TriviaQA | 73.32 | 74.10 |
| BBH | 70.65 | 69.57 |
| DROP | 71.37 | 68.17 |
| GSM8K | 77.63 | 76.65 |
| Repo Code PPL | 1.1351 | 1.1353 |
| RULER | 96.32 | 96.01 |
| LongPPL | 3.6053 | 3.4202 |
KV Bridging produces broadly comparable quality, with improvements on MMLU, TriviaQA, and LongPPL but declines on BBH, DROP, and GSM8K. In a separate comparison with KV Mirror, the final RULER scores are 87.65 for KV Bridging and 81.28 for KV Mirror.
Limitations
HySparse2 retains several costs and risks:
- Selection errors: relevant tokens not selected by full-attention scores are inaccessible to sparse layers.
- Residual full-attention cost: five full-attention layers remain in the full model, and one self-decoder full-attention layer is still used during prefill.
- Irregular token access: token-level sparsity can be less hardware-friendly than regular block sparsity.
- Budget competition: forced recent-token inclusion consumes part of the sparse support available for globally retrieved tokens.
- Task-dependent local modeling: Forced SWA is not uniformly superior to a separate gated SWA branch.
- Representational restriction: bridged K/V states are projections of self-decoder full-attention states rather than cross-decoder-native hidden states.
- Confounded comparisons: HySparse2 uses MQA while HySparse and Hybrid SWA use GQA, so cache and efficiency differences do not arise solely from the sparsity pattern.
- Mixed general reasoning: long-context retrieval improves substantially, but some general reasoning metrics decline.
- Incomplete end-to-end throughput evidence: the supplied results report cache sizes and prefill FLOP reductions, but not a comprehensive decoding latency or tokens-per-second table.
HySparse2 is therefore best characterized as a coordinated architecture for long-context agentic inference rather than merely a sparse-attention kernel. Its defining pipeline is
7
The central systems consequence is that prefill can terminate after the self-decoder, while the central modeling consequence is that sparse layers retain both recent local context and globally selected individual tokens.