Extrapolating recurrence depth at test time

Develop training and architectural methods for depth-recurrent transformer language models that enable reliable extrapolation to greater recurrence depths at test time, allowing the models to solve problems that are harder than those encountered during training while maintaining stability and performance.

Background

The paper proposes converting pretrained non-recurrent LLMs into depth-recurrent models via continued pretraining and shows benefits for math reasoning under fixed training compute. Depth-recurrence allows increasing test-time compute by iterating a recurrent block more times without increasing parameter count.

While retrofitting recurrence and scheduling strategies improve training efficiency and test-time gains, the authors highlight that building depth-recurrent models which can effectively recur deeper at inference than during training remains unresolved. This capability would enable solving harder problems by scaling internal latent computation beyond training settings.

References

One unsolved problem is how to most effectively build depth-recurrent models that can recur deeper at test time to solve harder problems than were seen during training.

— Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence  (2511.07384 - McLeish et al., 10 Nov 2025) in Discussion, Section 5

In principle, this could delay the eventual accuracy collapse observed in BUT. The present experiments do not test whether it does so.

— LSTM-UT and Recurrent-Depth Transformers on Cellular Automata  (2609.19521 - Kavuncu, 17 Sep 2026) in Section 7.5, “Interpretation and Limitations”; reiterated in Section 8, “Conclusion”

Finally, models with different numbers of thought tokens are trained separately; generalization to more thought tokens or adaptive thought-token counts at inference remains an open question.

— Trading Depth for Time in Recurrent Transformers  (2609.21605 - Huang et al., 18 Sep 2026) in Section Discussion and Limitations

All reported checkpoints use fixed loop schedules; additional test-time recurrence has not been evaluated.

— LoopVAE: Recurrent Depth Across Scales for Visual Tokenization  (2609.11516 - Lu, 10 Sep 2026) in Section 2, subsection “Scale-Conditioned Depth Recurrence”