Papers
Topics
Authors
Recent
Search
2000 character limit reached

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers

Published 16 Jul 2026 in cs.LG and cs.CL | (2607.15456v1)

Abstract: Looped, weight-tied Transformers reduce parameters by reusing a block, but decoding still stores a separate K/V cache for every recurrence step. We show that this loop-indexed cache is highly structured. For a fixed token, layer and head, K/V vectors trace a short low-rank trajectory across loops, while the head and layer axes remain much flatter. We introduce Looped Latent Attention (LLA), a post-training cache codec that stores compact K and V latents and reconstructs loop-specific K/V vectors only when attention reads them. The default per-head codec compresses recurrence, while LLA-2D also folds heads into one latent for the extreme-compression regime. The codec is initialized from SVD of teacher activations and refined with KL and attention-output distillation. At matched cache budget, per-head LLA outperforms head-axis MLA, cross-layer sharing, KV quantization and final-loop reuse, showing that the recurrent cache is low-rank but not safely collapsible to a single state. The same axis advantage holds on Ouro-2.6B-Thinking and transfers to Huginn-3.5B, where an SVD codec remains near-lossless to 32x in decoder-independent evaluation. The cache reduction is exact. On one H200, the latent-store path increases measured Ouro-1.4B batch capacity at 4k context from 32 to 768 sequences at 21.3x compression. For long math rollouts, on-policy refinement on student-generated prefixes raises MATH-500 at 4x from 0.43 to 0.66 and reduces no-answer generations.

Authors (2)

Summary

  • The paper introduces a post-training low-rank codec that compresses the recurrence axis of key/value caches, achieving exact ratios up to 21.3× while preserving near-teacher GSM8K accuracy across Ouro and Huginn models.
  • Cross-loop cache trajectories have substantially lower effective rank than head or layer axes, and asymmetric value/key ranks outperform shared-rank compression at the same memory budget.
  • On-policy refinement reduces long-generation errors, while deployment tests show up to 24× larger batches and major memory savings, although latent reconstruction can remain compute-bound and extreme compression harms retrieval.

Looped, weight-tied Transformers such as Ouro/LoopLM and Huginn reduce parameter count by applying a single block repeatedly, but at inference they still materialize a separate K/V cache entry for every recurrence step. A model with T=4T=4 loops over D=24D=24 layers therefore carries the KV footprint of a 96-layer decoder while sharing 24 layers of parameters. "Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers" (2607.15456) identifies this recurrence axis as the most compressible dimension of the looped-model cache and introduces a post-training codec that exploits it.

Method

For each layer, head and axis (K or V), LLA stacks the TT per-loop K/V vectors into one vector and encodes it as a centered low-rank latent via a down-projection WdownW_{\rm down}; each loop is reconstructed by a loop-specific up-projection plus per-loop mean. The codec stores (rk+rv)H(r_k+r_v)H scalars per token per layer instead of 2THdhead2THd_{\rm head}, giving an exact compression ratio ρ=2Tdhead/(rk+rv)\rho = 2Td_{\rm head}/(r_k+r_v). K and V use separate ranks with rv>rkr_v > r_k, because the value trajectory stabilizes more slowly across loops than the key trajectory — a matched-budget ablation confirms that the asymmetric split beats a shared-rank code at equal scalar count.

The codec is initialized from the top-rr right singular vectors of teacher activations collected on a small calibration set and fine-tuned with the teacher frozen, using forward KL distillation plus an attention-output matching term (λattn=0.5\lambda_{\rm attn}=0.5). SVD initialization dominates every other design choice: after a short conversion it reaches held-out KL of 0.105 versus 0.771 for random initialization at D=24D=240. Because a token occupies the same position at every recurrence step, RoPE phase is shared across loops and can be applied after reconstruction, avoiding MLA's decoupled-RoPE branch on the content path. A family variant, LLA-2D, additionally folds the head axis into a single latent per token and layer, trading accuracy for a further factor-D=24D=241 cache reduction in the extreme-compression regime.

Cross-loop structure in the frozen teacher

The paper's mechanistic claim is established directly on the cache tensor before any training. In both Ouro-1.4B and Ouro-2.6B-Thinking, the loop axis has normalized effective rank around 0.61–0.63, with a single direction carrying roughly 60% of the energy, while head and layer axes are near full rank by the same measure. Crucially, the trajectories are low-rank but not collapsed: only about 6% of K variance and 17% of V variance lie on the dominant direction, so the correct object is a short low-rank path rather than a fixed point. The same pattern transfers to Huginn-3.5B at D=24D=242 recurrence steps, supporting the interpretation that low cross-loop rank is a property of weight-tied iteration itself rather than an artifact of one model family.

Iso-cache results

The central comparison holds the KV-cache budget fixed across compression axes. At D=24D=243, per-head LLA achieves GSM8K 0.800 versus 0.522 for head-axis MLA, 0.660 for cross-layer sharing, and 0.702 for KV quantization, with train-KL roughly half or less of the learned baselines. The most striking control is final-loop reuse: a zero-parameter baseline with the identical D=24D=244 budget collapses GSM8K to 0.000, providing direct evidence that the loop trajectory cannot be replaced by its endpoint despite its apparent convergence.

The ordering survives scale and family changes. On Ouro-2.6B-Thinking, LLA remains near-lossless on GSM8K from D=24D=245 to D=24D=246 (0.834 to 0.821 against a teacher near 0.83), while the head baseline loses up to 0.25 absolute accuracy at D=24D=247 and quantization collapses to 0.351 at D=24D=248. On Huginn-3.5B, a training-free SVD codec is near-lossless to D=24D=249 in decoder-independent fidelity, and KL distillation recovers much of the aggressive TT0–TT1 range; head-axis reconstruction is 20–30× worse at matched budget, and LLA-2D becomes parameter-infeasible there because its input dimension grows with loop count.

On Ouro-1.4B, light compression behaves as a drop-in replacement: at TT2 the codec matches or slightly exceeds the teacher across GSM8K, HumanEval, MBPP, BBH and MMLU-Pro. Accuracy plateaus around 0.575–0.59 from TT3 to TT4 before rising sharply at TT5, indicating a rank threshold below which decision-critical tokens are not resolved. MATH-500 is consistently the most sensitive task, which motivates the next section.

Long-generation stability

Teacher-forced conversion trains on prefixes the deployed compressed model never visits; early reconstruction errors compound autoregressively over long rollouts. Binning completions by length shows off-policy accuracy falling well below the teacher with a spiking no-answer rate in the longest bin, while the teacher's own mild decline reflects problem difficulty only. The fix is an on-policy refinement stage: sample rollouts from the current codec (no reward filtering), apply a stop-gradient through sampling, and minimize the frozen teacher's per-token KL on those student-generated prefixes. This raises MATH-500 by +0.16 to +0.24 absolute at TT6–TT7 (e.g., 0.43 → 0.66 at TT8) and reduces no-answer generations at every ratio. At TT9 it improves termination but not accuracy, indicating the smallest latents have already lost task-relevant information — a limitation the authors state plainly.

Serving capacity and composition

The scalar reduction is exact and translates into measured capacity. On a single H200 at 4096-token context, maximum batch for Ouro-1.4B rises from 32 (teacher) to 128 at WdownW_{\rm down}0 and 768 at WdownW_{\rm down}1; the 1M-context cache falls from 768 GiB to 36 GiB when combined with token eviction. Eviction composition is clean: StreamingLLM/H2O-style eviction alone scores 0.510 on GSM8K under a 260-token budget, and stacking it on LLA-WdownW_{\rm down}2 scores 0.520, confirming orthogonality of the sequence and recurrence axes. Passkey retrieval is perfect through WdownW_{\rm down}3 up to 64k context, degrades at WdownW_{\rm down}4 beyond 16k, and fails entirely at WdownW_{\rm down}5, delimiting where exact token identity is lost.

Two implementation regimes are distinguished honestly. A full-memory fast path caches reconstructed K/V at 1.16× teacher latency but realizes no memory saving; a latent-store capacity path realizes the WdownW_{\rm down}6 saving but is reconstruction-compute-bound (108 vs 748 tok/s peak throughput). An absorbed attention implementation (moving K up-projections to the query side, MLA-style decoupled RoPE) measures 2.3× faster decode than reconstruct-then-attend at 262k context on a random-weight model, but post-hoc decoupled RoPE has a fidelity floor because the frozen queries were trained for full-RoPE keys; the authors treat absorption as a pretraining-native opportunity, not part of the main accuracy claim.

Limitations and open questions

The evaluation relies on small subsets for several generative benchmarks (e.g., 100–200 items for GSM8K codecs, 135 docs for BBH), so fine-grained differences within regimes carry sampling noise. The capacity path's compute-bound decoding means memory savings do not currently translate into latency wins, and the absorbed path's fidelity floor leaves open whether post-training conversion can ever match full-RoPE fidelity without native retraining. Whether the residual-stream update across loops is itself low enough rank to support computing shared content plus cheap per-loop deltas — an architecture change rather than a codec — remains untested. Finally, the sharp rank transition on GSM8K suggests but does not explain which token categories require latent resolution above the threshold.

Conclusion

This paper establishes the recurrence axis as the spectrally steepest and practically best target for KV-cache compression in weight-tied looped Transformers, supported by frozen-cache spectra, a decisive final-loop-reuse control, matched-budget baselines, scale generalization to Ouro-2.6B, and family transfer to Huginn-3.5B. On-policy refinement addresses the exposure-bias failure specific to long rollouts, and the exact memory reduction yields measured serving gains up to 24× batch capacity. The result reframes serving cost for looped models: as reasoning depth grows through recurrence, it is the cache, not the parameters, that must be compressed.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.