Assess whether Muon can be systematically improved
Ascertain whether the Muon optimizer can be systematically improved beyond its current formulation that orthogonalizes matrix-shaped gradient updates via spectral flattening.
References
Consequently, it remains unclear to what extent the reported improvements can be attributed to Muon itself, how Muon relates to established adaptive optimizers such as Adam, and whether Muon can be systematically improved.
Buffer-level compositions with Muon, Shampoo, and SOAP remain future work, and each deserves its own controlled study.
Future work could investigate shift-invariant formulations as well as other measures of distributional change, such as KL divergence, and evaluate whether their optimization benefits justify any additional computational cost.
However, persistent spectral constraints have not yet been shown to consistently achieve lower LLM pretraining loss than Muon.