Rigorous convergence rates for norm-based and preconditioned optimizers

Establish rigorous convergence rates for the optimization methods surveyed in this work—including architecture-aware preconditioners and norm-based optimizers such as KFAC, EKFAC, Shampoo, SOAP, SPlus, and Muon—on non-convex deep neural network objectives, and identify assumptions and step-size regimes under which these rates hold.

Background

The paper synthesizes classical and modern optimizers, emphasizing curvature-aware and norm-based designs (e.g., Kronecker-factored and spectral-geometry methods). While empirical successes are noted, formal convergence theory in the non-convex deep learning regime lags behind.

The authors explicitly call for rigorous convergence results tailored to these methods, reflecting a broader gap between practical performance and theoretical guarantees.

References

While this thesis has emphasized practical effectiveness and intuitive understanding, establishing rigorous convergence rates for the methods discussed — particularly in non-convex settings characteristic of deep learning — remains largely open.

— Towards Guided Descent: Optimization Algorithms for Training Neural Networks At Scale  (2512.18373 - Nagwekar, 20 Dec 2025) in Subsection “Theoretical Convergence Guarantees” within Section “Limitations of the Modular Norm Framework”

Several questions remain open. On the theoretical side, establishing conditions under which improved operator-oracle alignment translates into faster convergence remains an important next step.

— Muon-C: Operator-Aligned Muon for Convolutional Kernels  (2609.09676 - Qing et al., 9 Sep 2026) in Section Discussion

Consequently, we make no claim of unconditional convergence for this stochastic setting.

— AF-Muon: An AdamW-Free Muon Optimizer for Tied-Embedding Models  (2610.01395 - Lagzian et al., 1 Oct 2026) in Appendix, Section 'Exact oracle properties and conditional convergence bounds' (Appendix \ref{app:sec-convergenceAFMuon}); also stated in Section 2, subsection 'Full optimizer'

Extending the analysis to the full nonlinear and momentum-based updates of practical optimizers remains an open problem.