2000 character limit reached
A Note on the Convergence of Muon (2502.02900v2)
Published 5 Feb 2025 in math.OC
Abstract: In this note, we inspect the convergence of a new optimizer for pretraining LLMs, namely the Muon optimizer. Such an optimizer is closely related to a specialized steepest descent method where the update direction is the minimizer of the quadratic approximation of the objective function under spectral norm. We provide the convergence analysis on both versions of the optimizer and discuss its implications.
Collections
Sign up for free to add this paper to one or more collections.