Extrapolating recurrence depth at test time
Develop training and architectural methods for depth-recurrent transformer language models that enable reliable extrapolation to greater recurrence depths at test time, allowing the models to solve problems that are harder than those encountered during training while maintaining stability and performance.
References
One unsolved problem is how to most effectively build depth-recurrent models that can recur deeper at test time to solve harder problems than were seen during training.
In principle, this could delay the eventual accuracy collapse observed in BUT. The present experiments do not test whether it does so.
Finally, models with different numbers of thought tokens are trained separately; generalization to more thought tokens or adaptive thought-token counts at inference remains an open question.
All reported checkpoints use fixed loop schedules; additional test-time recurrence has not been evaluated.