Rigorous convergence rates for norm-based and preconditioned optimizers
Establish rigorous convergence rates for the optimization methods surveyed in this work—including architecture-aware preconditioners and norm-based optimizers such as KFAC, EKFAC, Shampoo, SOAP, SPlus, and Muon—on non-convex deep neural network objectives, and identify assumptions and step-size regimes under which these rates hold.
References
While this thesis has emphasized practical effectiveness and intuitive understanding, establishing rigorous convergence rates for the methods discussed — particularly in non-convex settings characteristic of deep learning — remains largely open.
Several questions remain open. On the theoretical side, establishing conditions under which improved operator-oracle alignment translates into faster convergence remains an important next step.
Consequently, we make no claim of unconditional convergence for this stochastic setting.
Extending the analysis to the full nonlinear and momentum-based updates of practical optimizers remains an open problem.