Papers
Topics
Authors
Recent
Search
2000 character limit reached

HySparse2 Long-Context Architectured: Hybrid with KV

Updated 24 September 2026
  • HySparse2 is a long-context transformer architecture designed for agentic inference, which leverages hybrid sparse attention with KV Bridging and KV Reuse mechanisms that optimize full-attention layers, and cross-decoder KV caches, ultimately improving efficiency by reducing cache storage and computational iteration in tasks requiring length and contextual integrity,.
  • HySparse2 addresses long-horizon, multi-turn agentic workloads by efficiently managing context through KV Bridging and KV Reuse. KV Bridging minimizes redundancy between the self- and cross-decoder. KV Reuse allows sparse layers to efficiently reuse key-value caches and token selections.
  • Token-level selection in HySparse2 improves sparse attention with a 128 recent local token forced window and 1,024 global selected tokens, aiding performance in benchmarks such as RULER-v2 and MRCR-v2.

HySparse2 is a long-context Transformer architecture for agentic inference that combines hybrid sparse attention with two-level key–value (KV) sharing (Wei et al., 22 Sep 2026). Its outer mechanism, KV Bridging, adopts a YOCO-style self-decoder/cross-decoder organization and constructs cross-decoder full-attention KV caches from self-decoder hidden states. Its inner mechanism, KV Reuse, shares full-attention KV caches and token-selection indices with subsequent sparse-attention layers. HySparse2 further replaces block-level sparsity with token-level selection and incorporates a forced recent-token window directly into sparse attention. The resulting architecture enables prefill to terminate after the self-decoder while preserving cross-decoder execution during autoregressive decoding.

1. Workload and architectural motivation

HySparse2 targets long-horizon, multi-turn agentic workloads in which short actions or tool calls are followed by long observations, such as search results, documents, tool outputs, execution traces, and intermediate reasoning. Each new observation expands the accumulated context, creating simultaneous demands for efficient prefill, compact KV-cache storage, and accurate long-context retrieval.

Conventional full attention has quadratic prefill cost with respect to sequence length. For queries QQ, keys KK, values VV, and causal mask MM, attention is represented schematically as

A(Q,K,V)=softmax⁡(QK⊤d+M)V.A(Q,K,V)=\operatorname{softmax}\left(\frac{QK^\top}{\sqrt d}+M\right)V.

The prefill score matrix contains O(L2)O(L^2) entries for sequence length LL. During decoding, attention to a cached history remains approximately O(Ld)O(Ld) per generated token, while storing K/V states for all layers grows linearly with LL but includes the multiplicative costs of layers, KV heads, head dimensions, and numerical precision.

HySparse2 does not remove global attention entirely. Instead, a small number of full-attention layers perform global mixing and provide attention scores for token selection. Subsequent sparse layers reuse the resulting KV representations and selected indices. This design allocates exact global computation to index-generating layers while limiting later attention operations to selected tokens.

The architecture is motivated particularly by contexts in which relevant evidence is distributed across multiple turns and may occur at arbitrary token positions. Token-level selection avoids retaining entire fixed-size blocks when only one or a few tokens in those blocks are relevant.

2. Two-level KV sharing

HySparse2 has two complementary levels of KV sharing.

Outer level—KV Bridging: the model is divided into a self-decoder and a cross-decoder. Full-attention layers in the self-decoder provide the hidden states from which the K/V caches of full-attention layers in the cross-decoder are generated.

Inner level—KV Reuse: within hybrid blocks, a full-attention layer constructs a full KV cache and token-selection indices. Subsequent sparse-attention layers reuse the selected KV entries and the corresponding indices instead of constructing independent full-length caches.

The combined structure is:

self-decoder hidden states→KV Bridgingcross-decoder full-attention KV caches\text{self-decoder hidden states} \xrightarrow{\text{KV Bridging}} \text{cross-decoder full-attention KV caches}

and

KK0

These mechanisms address different forms of redundancy. KV Bridging eliminates the need to generate cross-decoder KV caches from cross-decoder prefill hidden states. KV Reuse prevents multiple sparse layers following the same full-attention layer from independently materializing equivalent or overlapping long-context representations.

3. KV Bridging and prefill early exit

The 49-layer HySparse2 model is divided into a self-decoder and a cross-decoder. The self-decoder uses full attention and sliding-window attention (SWA), whereas the cross-decoder uses full attention and sparse attention.

For a self-decoder full-attention layer KK1 and a cross-decoder full-attention layer KK2, the bridged caches are

KK3

KK4

while the cross-decoder query is computed from its own hidden state:

KK5

Thus, KV Bridging shares the source representation for keys and values but does not share queries. Each cross-decoder full-attention layer has its own K/V projections, even when several layers use hidden states originating from the same self-decoder layer.

Only full-attention layers are bridged. The paper’s stated rationale is that full-attention hidden states incorporate global information and are therefore more suitable for generating cross-decoder KV caches than SWA states. One self-decoder full-attention layer may supply multiple cross-decoder full-attention layers, allowing the self- and cross-decoders to contain different proportions of full-attention layers.

During prefill, the execution is:

  1. The self-decoder processes the entire newly prefilling sequence.
  2. Self-decoder full-attention states are projected into cross-decoder full-attention K/V caches.
  3. Full-attention layers provide token-selection information for sparse layers.
  4. Sparse layers reuse the bridged caches and selection metadata.
  5. Prefill terminates without executing the cross-decoder layers over the entire sequence.

The cross-decoder remains active during autoregressive decoding. It is skipped only during prefill cache construction. In the evaluated 49-layer configuration, prefill requires the first 25 layers—the self-decoder and bridging projections—instead of all 49 layers.

A separate SWA branch in the cross-decoder would undermine this early exit because its K/V states would depend on cross-decoder hidden states, which themselves depend recursively on preceding cross-decoder layers. HySparse2 removes that dependency by eliminating the separate cross-decoder SWA branch.

4. Hybrid attention and token selection

The self-decoder combines full attention with SWA. For position KK6 and window width KK7, the local support is

KK8

SWA restricts attention to this set:

KK9

The evaluated configuration uses a 128-token window. Self-decoder SWA layers use partial RoPE with 64 rotary dimensions and base VV0. Full-attention and sparse-attention layers use NoPE.

The cross-decoder contains full-attention indexer layers and sparse-attention layers. For query VV1, sparse attention operates on a selected token set VV2:

VV3

The selection scores are the full-attention logits,

VV4

subject to the causal constraint VV5. Subsequent sparse layers reuse the full-attention layer’s selected indices and KV entries.

Token-level sparsity

HySparse selected 1,024 global tokens in 64-token blocks. HySparse2 selects individual tokens. In the evaluated configuration, 128 recent local tokens are forced into the support and 1,024 additional global tokens are selected by score:

VV6

VV7

VV8

Token-level selection permits the budget to be distributed across multiple turns without retaining up to 63 neighboring tokens for each selected token, as would occur with 64-token blocks.

The token-versus-block ablation reported the following results:

Metric Block selection Token selection
RULER-v2 49.56 56.13
MRCR-v2, two needles 12.94 21.08
GraphWalks 29.38 34.92
MMLU-Pro 35.74 36.97
NoLiMa 40.27 38.43
BBH 61.93 60.70

Token-level sparsity improved the principal long-context retrieval metrics but was not uniformly superior on every task.

Forced recent-token selection

HySparse2 removes the separate SWA branch from sparse layers. Instead, recent tokens are inserted directly into the same sparse support:

VV9

Local and globally selected tokens are therefore read from the same full-attention KV cache. There is no separate local K/V projection, second local cache, or gated fusion branch.

The local-attention ablation compared Gated SWA, No SWA, and Forced SWA:

Metric Gated SWA No SWA Forced SWA
RULER 88.19 84.55 89.84
RULER-v2 53.66 54.62 55.98
MRCR-v2 27.66 20.73 22.67
GraphWalks 35.39 36.48 37.13
LongPPL 6.8807 7.1307 6.9838
GSM8K 64.52 60.35 59.44

Forced SWA enabled the early-exit prefill structure and performed strongly on several retrieval metrics, but it underperformed Gated SWA on MRCR-v2 and GSM8K and had higher LongPPL.

5. Model configuration and inference workflow

The evaluated models are 80B-A3B mixture-of-experts models with:

  • 49 Transformer layers;
  • hidden size 2,048;
  • approximately 80B total parameters and 3B active parameters;
  • FP8 KV-cache storage;
  • simplified mHC residual mixing with the residual matrix fixed to identity.

The attention configurations are:

Model Full-attention layers Q/KV heads QK/V head dimension
Hybrid SWA 9 64/4 192/128
HySparse 5 64/4 192/128
HySparse2 5 64/1 256/256

HySparse2 uses multi-query attention (MQA), whereas the baselines use grouped-query attention (GQA). This difference contributes to the KV-cache reduction and affects the comparison independently of the attention sparsity pattern.

The principal HySparse2 settings are:

  • 128 forced local tokens;
  • 1,024 global tokens;
  • token-level selection;
  • self-decoder SWA;
  • cross-decoder sparse attention;
  • NoPE in full and sparse attention;
  • partial RoPE in self-decoder SWA;
  • sigmoid output gates in sparse/SWA layers;
  • learnable per-head sink biases.

The training setup uses approximately 500B pretraining tokens with 32k pretraining context and approximately 100B post-training tokens. Agentic data are added during post-training, and the context is extended to 256k. The compared models use the same data and schedules, the Muon optimizer, and a WSD schedule. Peak learning rates are MM0 for pretraining and MM1 for post-training.

For a prefilling sequence, the workflow is:

  1. Run the self-decoder using full attention and SWA.
  2. Extract hidden states from self-decoder full-attention layers.
  3. Project those states into cross-decoder full-attention K/V caches.
  4. Generate full-attention token-selection metadata.
  5. Add the forced recent-token window to each sparse support.
  6. Transfer the bridged caches and selection metadata to decoding.
  7. End prefill after the self-decoder and bridging operations.
  8. Run the cross-decoder normally during autoregressive decoding.

During decoding, cross-decoder queries are computed from cross-decoder hidden states. Full-attention layers use bridged K/V caches, while sparse layers access selected entries from those caches.

6. Evaluation, efficiency, and limitations

Quality results

Selected pretraining results are:

Task Hybrid SWA HySparse HySparse2
RULER 88.71 84.89 90.77
NoLiMa 30.13 40.27 49.76
Repo Code PPL 1.1578 1.1588 1.1570
BBH 60.14 61.93 64.29
MMLU-Pro 36.12 35.74 37.56
DROP 60.90 63.78 58.99
GSM8K 64.44 64.14 61.94

HySparse2 shows its strongest advantage on long-context evaluation but underperforms HySparse on DROP and GSM8K.

The post-training evaluation includes AgentPPL, LongPPL, MRCR-v2, RULER-v2, and GraphWalks. Relative to HySparse, HySparse2 improves mean MRCR-v2 by 11.30 percentage points and mean RULER-v2 by 19.81 points. Relative to Hybrid SWA, the corresponding improvements are 6.44 and 18.65 points.

At 256k context, RULER-v2 scores are:

  • HySparse2: 58.45;
  • HySparse: 32.61;
  • Hybrid SWA: 35.74.

The paper reports lower AgentPPL and LongPPL for HySparse2 at all evaluated lengths. AgentPPL increases with context length, whereas LongPPL decreases, reflecting the different effects of longer context on multi-turn retrieval and long-range-dependent tokens.

Prefill computation and KV storage

At one million tokens with FP8 KV storage, the reported cache sizes are:

Quantity Hybrid SWA HySparse HySparse2
KV cache 12.09 GB 6.72 GB 2.69 GB

Reported prefill FLOP reductions at one million tokens are 2.92 times relative to HySparse and 5.02 times relative to Hybrid SWA. The reduction results from:

  • terminating prefill after the self-decoder;
  • using only one self-decoder full-attention layer during prefill in the evaluated configuration;
  • sharing cross-decoder KV caches;
  • reusing full-attention KV entries in sparse layers;
  • removing the separate sparse-layer SWA branch;
  • using MQA instead of GQA;
  • selecting only a limited token support in sparse attention.

A conventional KV-storage estimate is proportional to

MM2

where MM3 is sequence length, MM4 is the number of KV heads, MM5 is the head dimension, and MM6 is bytes per element. HySparse2 reduces this quantity through MQA, cross-layer sharing, and sparse reuse.

KV Bridging ablations

A separate KV Bridging study at 290B-A8B scale used approximately 1.8T training tokens and 32k context:

Task Without bridging With bridging
MMLU 72.68 72.80
TriviaQA 73.32 74.10
BBH 70.65 69.57
DROP 71.37 68.17
GSM8K 77.63 76.65
Repo Code PPL 1.1351 1.1353
RULER 96.32 96.01
LongPPL 3.6053 3.4202

KV Bridging produces broadly comparable quality, with improvements on MMLU, TriviaQA, and LongPPL but declines on BBH, DROP, and GSM8K. In a separate comparison with KV Mirror, the final RULER scores are 87.65 for KV Bridging and 81.28 for KV Mirror.

Limitations

HySparse2 retains several costs and risks:

  • Selection errors: relevant tokens not selected by full-attention scores are inaccessible to sparse layers.
  • Residual full-attention cost: five full-attention layers remain in the full model, and one self-decoder full-attention layer is still used during prefill.
  • Irregular token access: token-level sparsity can be less hardware-friendly than regular block sparsity.
  • Budget competition: forced recent-token inclusion consumes part of the sparse support available for globally retrieved tokens.
  • Task-dependent local modeling: Forced SWA is not uniformly superior to a separate gated SWA branch.
  • Representational restriction: bridged K/V states are projections of self-decoder full-attention states rather than cross-decoder-native hidden states.
  • Confounded comparisons: HySparse2 uses MQA while HySparse and Hybrid SWA use GQA, so cache and efficiency differences do not arise solely from the sparsity pattern.
  • Mixed general reasoning: long-context retrieval improves substantially, but some general reasoning metrics decline.
  • Incomplete end-to-end throughput evidence: the supplied results report cache sizes and prefill FLOP reductions, but not a comprehensive decoding latency or tokens-per-second table.

HySparse2 is therefore best characterized as a coordinated architecture for long-context agentic inference rather than merely a sparse-attention kernel. Its defining pipeline is

MM7

The central systems consequence is that prefill can terminate after the self-decoder, while the central modeling consequence is that sparse layers retain both recent local context and globally selected individual tokens.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to HySparse2.