Explanation of the Long-Horizon Training-Loss Optimum Shift
Determine whether the lower training-loss learning-rate preference observed after approximately four passes over WikiText-103 is caused by data repetition rather than by a genuine change in the learning-rate optimum during long-horizon training.
References
At 60k, we observe that the training loss favors $\eta{\leq}0.7$ while the validation loss still selects $\eta{=}1.0$; we conjecture that this reflects these runs passing over WikiText-103 about four times, and read the optimum from validation loss.
— Learning Rate Transfer for Hybrid Transformer-SSM Architectures
(2610.01172 - Seo et al., 1 Oct 2026) in Appendix, Section “Width 2048 and Billion Scale,” subsection “Billion scale, horizon, and seeds”