- The paper introduces a post-training low-rank codec that compresses the recurrence axis of key/value caches, achieving exact ratios up to 21.3× while preserving near-teacher GSM8K accuracy across Ouro and Huginn models.
- Cross-loop cache trajectories have substantially lower effective rank than head or layer axes, and asymmetric value/key ranks outperform shared-rank compression at the same memory budget.
- On-policy refinement reduces long-generation errors, while deployment tests show up to 24× larger batches and major memory savings, although latent reconstruction can remain compute-bound and extreme compression harms retrieval.
Looped, weight-tied Transformers such as Ouro/LoopLM and Huginn reduce parameter count by applying a single block repeatedly, but at inference they still materialize a separate K/V cache entry for every recurrence step. A model with T=4 loops over D=24 layers therefore carries the KV footprint of a 96-layer decoder while sharing 24 layers of parameters. "Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers" (2607.15456) identifies this recurrence axis as the most compressible dimension of the looped-model cache and introduces a post-training codec that exploits it.
Method
For each layer, head and axis (K or V), LLA stacks the T per-loop K/V vectors into one vector and encodes it as a centered low-rank latent via a down-projection Wdown; each loop is reconstructed by a loop-specific up-projection plus per-loop mean. The codec stores (rk+rv)H scalars per token per layer instead of 2THdhead, giving an exact compression ratio ρ=2Tdhead/(rk+rv). K and V use separate ranks with rv>rk, because the value trajectory stabilizes more slowly across loops than the key trajectory — a matched-budget ablation confirms that the asymmetric split beats a shared-rank code at equal scalar count.
The codec is initialized from the top-r right singular vectors of teacher activations collected on a small calibration set and fine-tuned with the teacher frozen, using forward KL distillation plus an attention-output matching term (λattn=0.5). SVD initialization dominates every other design choice: after a short conversion it reaches held-out KL of 0.105 versus 0.771 for random initialization at D=240. Because a token occupies the same position at every recurrence step, RoPE phase is shared across loops and can be applied after reconstruction, avoiding MLA's decoupled-RoPE branch on the content path. A family variant, LLA-2D, additionally folds the head axis into a single latent per token and layer, trading accuracy for a further factor-D=241 cache reduction in the extreme-compression regime.
Cross-loop structure in the frozen teacher
The paper's mechanistic claim is established directly on the cache tensor before any training. In both Ouro-1.4B and Ouro-2.6B-Thinking, the loop axis has normalized effective rank around 0.61–0.63, with a single direction carrying roughly 60% of the energy, while head and layer axes are near full rank by the same measure. Crucially, the trajectories are low-rank but not collapsed: only about 6% of K variance and 17% of V variance lie on the dominant direction, so the correct object is a short low-rank path rather than a fixed point. The same pattern transfers to Huginn-3.5B at D=242 recurrence steps, supporting the interpretation that low cross-loop rank is a property of weight-tied iteration itself rather than an artifact of one model family.
Iso-cache results
The central comparison holds the KV-cache budget fixed across compression axes. At D=243, per-head LLA achieves GSM8K 0.800 versus 0.522 for head-axis MLA, 0.660 for cross-layer sharing, and 0.702 for KV quantization, with train-KL roughly half or less of the learned baselines. The most striking control is final-loop reuse: a zero-parameter baseline with the identical D=244 budget collapses GSM8K to 0.000, providing direct evidence that the loop trajectory cannot be replaced by its endpoint despite its apparent convergence.
The ordering survives scale and family changes. On Ouro-2.6B-Thinking, LLA remains near-lossless on GSM8K from D=245 to D=246 (0.834 to 0.821 against a teacher near 0.83), while the head baseline loses up to 0.25 absolute accuracy at D=247 and quantization collapses to 0.351 at D=248. On Huginn-3.5B, a training-free SVD codec is near-lossless to D=249 in decoder-independent fidelity, and KL distillation recovers much of the aggressive T0–T1 range; head-axis reconstruction is 20–30× worse at matched budget, and LLA-2D becomes parameter-infeasible there because its input dimension grows with loop count.
On Ouro-1.4B, light compression behaves as a drop-in replacement: at T2 the codec matches or slightly exceeds the teacher across GSM8K, HumanEval, MBPP, BBH and MMLU-Pro. Accuracy plateaus around 0.575–0.59 from T3 to T4 before rising sharply at T5, indicating a rank threshold below which decision-critical tokens are not resolved. MATH-500 is consistently the most sensitive task, which motivates the next section.
Long-generation stability
Teacher-forced conversion trains on prefixes the deployed compressed model never visits; early reconstruction errors compound autoregressively over long rollouts. Binning completions by length shows off-policy accuracy falling well below the teacher with a spiking no-answer rate in the longest bin, while the teacher's own mild decline reflects problem difficulty only. The fix is an on-policy refinement stage: sample rollouts from the current codec (no reward filtering), apply a stop-gradient through sampling, and minimize the frozen teacher's per-token KL on those student-generated prefixes. This raises MATH-500 by +0.16 to +0.24 absolute at T6–T7 (e.g., 0.43 → 0.66 at T8) and reduces no-answer generations at every ratio. At T9 it improves termination but not accuracy, indicating the smallest latents have already lost task-relevant information — a limitation the authors state plainly.
Serving capacity and composition
The scalar reduction is exact and translates into measured capacity. On a single H200 at 4096-token context, maximum batch for Ouro-1.4B rises from 32 (teacher) to 128 at Wdown0 and 768 at Wdown1; the 1M-context cache falls from 768 GiB to 36 GiB when combined with token eviction. Eviction composition is clean: StreamingLLM/H2O-style eviction alone scores 0.510 on GSM8K under a 260-token budget, and stacking it on LLA-Wdown2 scores 0.520, confirming orthogonality of the sequence and recurrence axes. Passkey retrieval is perfect through Wdown3 up to 64k context, degrades at Wdown4 beyond 16k, and fails entirely at Wdown5, delimiting where exact token identity is lost.
Two implementation regimes are distinguished honestly. A full-memory fast path caches reconstructed K/V at 1.16× teacher latency but realizes no memory saving; a latent-store capacity path realizes the Wdown6 saving but is reconstruction-compute-bound (108 vs 748 tok/s peak throughput). An absorbed attention implementation (moving K up-projections to the query side, MLA-style decoupled RoPE) measures 2.3× faster decode than reconstruct-then-attend at 262k context on a random-weight model, but post-hoc decoupled RoPE has a fidelity floor because the frozen queries were trained for full-RoPE keys; the authors treat absorption as a pretraining-native opportunity, not part of the main accuracy claim.
Limitations and open questions
The evaluation relies on small subsets for several generative benchmarks (e.g., 100–200 items for GSM8K codecs, 135 docs for BBH), so fine-grained differences within regimes carry sampling noise. The capacity path's compute-bound decoding means memory savings do not currently translate into latency wins, and the absorbed path's fidelity floor leaves open whether post-training conversion can ever match full-RoPE fidelity without native retraining. Whether the residual-stream update across loops is itself low enough rank to support computing shared content plus cheap per-loop deltas — an architecture change rather than a codec — remains untested. Finally, the sharp rank transition on GSM8K suggests but does not explain which token categories require latent resolution above the threshold.
Conclusion
This paper establishes the recurrence axis as the spectrally steepest and practically best target for KV-cache compression in weight-tied looped Transformers, supported by frozen-cache spectra, a decisive final-loop-reuse control, matched-budget baselines, scale generalization to Ouro-2.6B, and family transfer to Huginn-3.5B. On-policy refinement addresses the exposure-bias failure specific to long rollouts, and the exact memory reduction yields measured serving gains up to 24× batch capacity. The result reframes serving cost for looped models: as reasoning depth grows through recurrence, it is the cache, not the parameters, that must be compressed.