Explanation of the Long-Horizon Training-Loss Optimum Shift

Determine whether the lower training-loss learning-rate preference observed after approximately four passes over WikiText-103 is caused by data repetition rather than by a genuine change in the learning-rate optimum during long-horizon training.

Background

At 1.24 billion parameters and 60,000 training steps, the validation loss still selects learning rate 1.0, whereas the training loss favors learning rates at or below 0.7. The authors hypothesize that repeated passes over WikiText-103 explain this discrepancy, but they do not establish the cause.

References

At 60k, we observe that the training loss favors $\eta{\leq}0.7$ while the validation loss still selects $\eta{=}1.0$; we conjecture that this reflects these runs passing over WikiText-103 about four times, and read the optimum from validation loss.

— Learning Rate Transfer for Hybrid Transformer-SSM Architectures  (2610.01172 - Seo et al., 1 Oct 2026) in Appendix, Section “Width 2048 and Billion Scale,” subsection “Billion scale, horizon, and seeds”