MuonIO: Principled Norm-Aware Descent for Embedding Tables and Language Model Heads
Abstract: The Muon optimizer derives its update rule for hidden linear layers by solving a local linearization of the loss penalized by the spectral norm, motivated by an RMS-stability argument for dense linear layers. Standard Muon implementations, however, exclude the input (embedding table) and output (LLM head) layers from this principled treatment, for which they use AdamW instead. We present MuonIO, a single Muon-style update for both of these layers. For the LLM head , we motivate the use of the operator norm, due to the Lipschitz continuity of the softmax output geometry, while for the embedding table , we draw on the operator norm, based on the one-hot input geometry identified by Bernstein & Newhouse (2025). The identity then puts both matrices in the same vocabulary-oriented geometry: MuonIO applies a single normalized-vector rule, which appears as column normalization for and row normalization for . Empirical evaluations demonstrate the effectiveness of our approach, with MuonIO reducing I/O optimizer state memory by 50% and I/O update FLOPs by 46% compared to Muon for 1B LLaMA pretraining on C4, while also improving validation perplexity.
Paper Prompts
Sign up for free to create and run prompts on this paper.