Adaptive-halting evaluation at 4B scale
Conduct a full adaptive-halting evaluation of RecurTrace at the 4B model scale to determine whether its learned halting mechanism retains the reported benefits beyond the fixed two-loop scaling experiment.
References
The halting head is distilled at $4$B as at the other scales (Stage~2, Table~\ref{tab:budget}); we report this scaling row at a fixed two-loop depth and leave a full adaptive-halting evaluation at $4$B to future work.
— RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory
(2609.03379 - Wang et al., 3 Sep 2026) in Appendix, Section "Additional Ablations and Discussion," paragraph "The 4B Comparison"