Ultimate extrapolation boundaries of Fitting and Transfer paradigms
Ascertain the maximum scale at which learning-rate predictions derived from the Fitting paradigm and from μTransfer-based hyperparameter transfer remain accurate in large-scale pre-training, thereby identifying the ultimate extrapolation boundaries for these approaches.
References
Due to computational resource constraints, this study did not investigate the ultimate extrapolation boundaries (i.e., the maximum scale at which these predictions remain accurate) for both the Fitting and Transfer paradigms.
Billion scale, the full training horizon, and multiple seeds are validated individually rather than jointly: the LR-transfer sweep reaches 1.24B parameters (three seeds, 20k steps), and one 1.24B model trains cleanly to a Chinchilla-optimal horizon (24.9B tokens) at the transferred LR with a single LR and seed (Appendix~\ref{sec:appendix:scale}); whether the near-zero gap persists when all three hold at once remains open.