Assess whether Muon can be systematically improved

Ascertain whether the Muon optimizer can be systematically improved beyond its current formulation that orthogonalizes matrix-shaped gradient updates via spectral flattening.

Background

In motivating their study, the authors raise the question of whether Muon can be enhanced in a structured way. They later propose a spectral family of transformations and fractional variants to explore potential improvements, but the general question of systematic improvement is posed explicitly.

References

Consequently, it remains unclear to what extent the reported improvements can be attributed to Muon itself, how Muon relates to established adaptive optimizers such as Adam, and whether Muon can be systematically improved.

— Delving into Muon and Beyond: Deep Analysis and Extensions  (2602.04669 - Qi et al., 4 Feb 2026) in Section 1 (Introduction)

Buffer-level compositions with Muon, Shampoo, and SOAP remain future work, and each deserves its own controlled study.

— DeltaMomentum: A Key-Value based Anisotropic Momentum Update via Delta Rule  (2608.19491 - Hong et al., 19 Aug 2026) in Section 7, paragraph "Limitations and Future Work"

Future work could investigate shift-invariant formulations as well as other measures of distributional change, such as KL divergence, and evaluate whether their optimization benefits justify any additional computational cost.

— MuonIO: Principled Norm-Aware Descent for Embedding Tables and Language Model Heads  (2610.02705 - Ma et al., 2 Oct 2026) in Section 6, “Limitations and future work”

However, persistent spectral constraints have not yet been shown to consistently achieve lower LLM pretraining loss than Muon.

— ORCA: The Annealed Spectral Conditioning Optimizer for Faster, Better LLM Training  (2610.06116 - Liu et al., 5 Oct 2026) in Appendix, Section “Orthogonal and spectral training”