Decoding Looped Transformers Better for (Almost) Free
This presentation explores a breakthrough in looped Transformer inference. Looped architectures recycle Transformer blocks to achieve greater depth without adding parameters, but conventional decoding discards all intermediate recurrent states except the last. The authors show that these earlier states are not computational waste—they are naturally aligned weak predictors that can guide the final output through contrastive decoding. The resulting method, LoopCD, requires no training, no auxiliary models, and adds negligible computational cost while delivering substantial accuracy gains on mathematical reasoning, code generation, and multiple-choice tasks. Most strikingly, using only half the recurrent iterations with LoopCD matches or exceeds full-depth unguided accuracy while cutting forward compute by up to 48%.Script
Looped Transformers run the same block over and over to get deeper predictions without adding parameters, but every architecture throws away all the intermediate states except the very last one. This paper proves those discarded states are not waste: they are weak predictors that can guide the final answer, delivering better accuracy for almost no extra cost.
The method, called LoopCD, contrasts an early recurrent state with the final one. The early state is weaker because it has undergone fewer iterations, but it shares the same vocabulary, input context, and parameters. That alignment is the key: you get a weak model without training one, without loading a checkpoint, and without changing the prompt.
Two variants exist. LoopCD Logits contrasts output distributions after decoding both states independently, which requires two passes through the post-loop layers. LoopCD Hidden extrapolates the hidden states before the coda, then decodes once. The hidden form adds effectively zero output cost, and when the coda is shallow it can actually outperform the logit form.
On mathematical reasoning with the Ouro model, adaptive LoopCD Logits pushes AIME 2024 pass at 1 from 61.88% to 73.33%, an 11 point leap. AIME 2025 climbs 7 points, OlympiadBench over 3 points. For code generation, Huginn gains 9 points on HumanEval using the hidden-state form, and multiple-choice benchmarks show smaller but consistent cross-architecture improvements.
The systems payoff is even sharper. Cutting recurrent iterations in half normally drops multiple-choice accuracy by up to 1.3 points. Add LoopCD at the reduced depth and it not only recovers the loss but exceeds the full-depth unguided baseline, while theoretical forward operations drop 22% to 48% depending on architecture. You get better predictions with less compute.
Why does it work? The first recurrent state is weaker by 4 to 24 points, and its prediction diverges substantially from the final one. As iterations proceed, disagreement falls and standalone accuracy rises. LoopCD gains track that disagreement directly: when the early state picks a different answer, contrasting it re-ranks uncertain decisions without disrupting confident ones. The improvement is targeted correction, not global reshaping. You can learn more about contrastive decoding for looped architectures and create your own research videos at EmergentMind.com.