Adaptive-halting evaluation at 4B scale

Conduct a full adaptive-halting evaluation of RecurTrace at the 4B model scale to determine whether its learned halting mechanism retains the reported benefits beyond the fixed two-loop scaling experiment.

Background

RecurTrace is evaluated at multiple Qwen3 model scales, but the 4B scaling result reported in the paper uses a fixed two-loop depth rather than invoking the adaptive halting head. The 4B loop-module training run is also shorter than the schedules used at smaller scales, so the authors characterize this result as an initial scale check rather than a fully matched replication.

The paper does report halting-head distillation at 4B, but it does not present a corresponding full adaptive-halting benchmark at that scale. Establishing such an evaluation would test whether the method's compute-allocation behavior and accuracy gains generalize from the controlled 1.7B study to the larger 4B model.

References

The halting head is distilled at $4$B as at the other scales (Stage~2, Table~\ref{tab:budget}); we report this scaling row at a fixed two-loop depth and leave a full adaptive-halting evaluation at $4$B to future work.

RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory  (2609.03379 - Wang et al., 3 Sep 2026) in Appendix, Section "Additional Ablations and Discussion," paragraph "The 4B Comparison"