Papers
Topics
Authors
Recent
Search
2000 character limit reached

MuonIO: Principled Norm-Aware Descent for Embedding Tables and Language Model Heads

Published 2 Oct 2026 in cs.LG and cs.AI | (2610.02705v1)

Abstract: The Muon optimizer derives its update rule for hidden linear layers by solving a local linearization of the loss penalized by the spectral norm, motivated by an RMS-stability argument for dense linear layers. Standard Muon implementations, however, exclude the input (embedding table) and output (LLM head) layers from this principled treatment, for which they use AdamW instead. We present MuonIO, a single Muon-style update for both of these layers. For the LLM head L∈R<sup>V</sup>×d\mathbf{L} \in \mathbb{R}<sup>{V</sup> \times d}, we motivate the use of the 2→∞2\to\infty operator norm, due to the Lipschitz continuity of the softmax output geometry, while for the embedding table E∈R<sup>d</sup>×V\mathbf{E} \in \mathbb{R}<sup>{d</sup> \times V}, we draw on the 1→21 \to 2 operator norm, based on the one-hot input geometry identified by Bernstein & Newhouse (2025). The identity ∥L∥2→∞=∥L<sup>⊤∥1→2\lVert\mathbf{L}\rVert_{2\to\infty}=\lVert\mathbf{L}<sup>\top\rVert_{1\to2} then puts both matrices in the same vocabulary-oriented geometry: MuonIO applies a single normalized-vector rule, which appears as column normalization for E\mathbf{E} and row normalization for L\mathbf{L}. Empirical evaluations demonstrate the effectiveness of our approach, with MuonIO reducing I/O optimizer state memory by 50% and I/O update FLOPs by ∼\sim46% compared to Muon for 1B LLaMA pretraining on C4, while also improving validation perplexity.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.