Reliable hyperparameter transfer across scales
Develop reliable methods to transfer optimal training hyperparameters such as learning rate and initialization from small-scale proxy models to large-scale Large Language Models while guaranteeing stable training dynamics.
References
Consequently, a critical open question is how to reliably transfer optimal hyperparameters (e.g., learning rate, initialization) found on small-scale proxy models to large-scale target models.
Nevertheless, sparsity follows distinct scaling behaviors \citep{clark2022unified,tian2025towards,team2025kimi}, and it remains unclear whether existing hyperparameter transfer frameworks derived for width scaling generalize to the sparsity dimension. This critical gap remains largely unaddressed, leaving optimal hyperparameter configuration prediction for large-scale MoE training an open problem.