Papers
Topics
Authors
Recent
Search
2000 character limit reached

Muon Learns Facts Better: Understanding the Role of Spectral Orthogonalization

Published 2 Oct 2026 in cs.LG, math.OC, and stat.ML | (2610.02798v1)

Abstract: The Muon optimizer applies spectral orthogonalization to matrix-valued updates and has shown strong performance in large-scale neural network training, yet the mechanisms of this transformation in feature learning remain poorly understood. In this work, we investigate this question through a tractable factual-recall model, where a fact maps each subject-relation pair to an answer, and a linear transformer learns the subject- and relation-dependent information required to recover this mapping. The transformer is optimized with gradient flow (GF), spectral GF, or Sign GF, which are continuous-time limits of gradient descent, Muon, and Adam, respectively. Prior studies (Nichani et al., 2025) have shown that when the number of subjects exceeds the number of relations, GF learns relation-dependent information before subject-dependent information, producing a feature-separation phase during training. We characterize this separation with the learning times when the subject- and relation-dependent components of the prediction reach a target accuracy. With SS subjects and RR relations, GF has a learning-time ratio of Θ~(S/R)\widetildeΘ(\sqrt{S/R}), whereas Spectral GF reduces this ratio to Θ~(1)\widetildeΘ(1). In addition, for fixed SS and RR, the subject- and relation-dependent errors decay as 1/(Tlog⁡T)1/(T\log T) in training time TT under GF, but as exp⁡(−poly(T))\exp(-\mathrm{poly}(T)) under spectral GF. Finally, we show that GF and spectral GF are equivariant under orthogonal transformations of the token embeddings, whereas Sign GF is not: Different orthonormal embeddings can potentially produce no feature separation, a large feature-separation phase, or even a reversed learning order. These results provide a mechanistic view of how spectral orthogonalization can fundamentally reshape feature-learning dynamics.

Authors (4)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.