---
title: 'MuonH Optimizer: Stiefel Projection for ERM'
url: https://www.emergentmind.com/topics/muonh-optimizer
type: topic
---

# MuonH Optimizer: Stiefel Projection for ERM

MuonH is a variant of the Muon optimizer that incorporates orthonormal search directions via projection onto the Stiefel manifold, specifically designed for nonconvex empirical risk minimization (ERM) in the presence of heavy-tailed stochastic noise and Hölder-smoothness in the objective's gradients. MuonH generalizes normalized-SGD to the matrix setting and provides rigorous convergence guarantees under conditions where traditional assumptions, such as bounded variance, do not hold. The method leverages key advances in manifold optimization and heavy-tail resilient stochastic approximation to achieve faster convergence rates than standard mini-batch SGD in terms of the gradient norm, with particular relevance for deep neural networks and associative memory structures in large language models.

## 1. Algorithmic Structure of MuonH

MuonH operates on a matrix parameterization $W_t \in \mathbb{R}^{m \times n}$ at iteration $t$. Each update enforces orthogonality in the search direction by projecting a stochastic (mini-batch) gradient onto the Stiefel manifold. The central loop of MuonH (without momentum, $\beta=0$) is:

1. Draw a mini-batch $\xi_t = (\xi_{t,1},\dots,\xi_{t,b_t})$ from the dataset.
2. Compute the mini-batch gradient:
   $$
   G_t = \frac{1}{b_t}\sum_{i=1}^{b_t} \nabla f_{\xi_{t,i}}(W_t)
   $$
3. Project $G_t$ onto the Stiefel manifold to obtain the orthonormal direction $O_t$:
   $$
   O_t = \arg\min_{O \in \mathbb{R}^{m \times n} : O^\top O = I_n} \|O - G_t\|_F
   $$
   If $G_t = U_t \Sigma_t V_t^\top$ (compact SVD), $O_t = U_t V_t^\top$.
4. Update parameters:
   $$
   W_{t+1} = W_t - \eta_t O_t
   $$

Momentum can be incorporated by replacing $G_t$ with a weighted sum incorporating past gradients. Newton–Schulz iteration is used in practice as an efficient alternative to the full SVD for approximating $O_t$ [2603.15059].

## 2. Theoretical Foundations: Hölder-Smoothness and Heavy-Tailed Noise

MuonH is analyzed for empirical risk minimization objectives
$$
\min_{W \in \mathbb{R}^{m \times n}} f(W), \quad f(W) = \frac{1}{N} \sum_{i=1}^N f_i(W)
$$
with two structural properties:

- **Hölder-Smoothness**: The gradients of $f_i$ satisfy for all $W_1, W_2$ and exponent $\nu \in (0,1]$
  $$
  \|\nabla f_i(W_1) - \nabla f_i(W_2)\|_F \le L_i\|W_1-W_2\|_F^\nu
  $$
- **Heavy-Tailed Noise**: The stochastic gradient estimator $\nabla f_\xi(W)$ is unbiased, with $p$-variance bounded by $\sigma^p$ for $p \in (1,2]$:
  $$
  \mathbb{E}_\xi[\|\nabla f_\xi(W) - \nabla f(W)\|_F^p] \le \sigma^p
  $$
  The regime $p<2$ admits genuinely heavy-tailed noise, encountered in practical large-scale learning [2603.15059].

## 3. Convergence Guarantees and Rate Improvements

The main convergence theorem establishes that, under suitable step-size and mini-batch schedules:
- $\sum_{t=0}^\infty \eta_t = \infty$
- $\sum_{t=0}^\infty \eta_t^{1+\nu} < \infty$
- $\sum_{t=0}^\infty \frac{\eta_t}{b_t^{(p-1)/p}} < \infty$
the iterates $\{W_t\}$ satisfy almost surely
$$
\liminf_{t\to\infty} \|\nabla f(W_t)\|_F = 0,
$$
i.e. convergence to stationary points even under heavy-tailed noise. The step-size can be taken as $\eta_t = \eta_0/(t+1)^a$ with $a \in (1/(1+\nu),1)$.

**Convergence Rate Comparison:**
- **Mini-batch SGD** achieves $\min_t \mathbb{E}[\|\nabla f(W_t)\|_F^2] = O(T^{a-1})$ so $\min_t \mathbb{E}[\|\nabla f(W_t)\|_F] = O(T^{(a-1)/2})$
- **MuonH** achieves $\min_t \mathbb{E}[\|\nabla f(W_t)\|_F] = O(T^{a-1})$, i.e., MuonH improves the rate in the gradient norm metric.

If $a \approx 1$ (e.g., $a=0.51$ for $\nu=1$), SGD's best achievable rate is $O(1/\sqrt{T})$ (squared-norm), while MuonH attains $O(1/T)$ (norm) [2603.15059].

## 4. Comparative Structure: Muon, MuonH, and AdamW

The Muon family is fundamentally distinct from standard optimizers through its use of matrix manifold geometry and explicit spectral (operator-norm) constraints:
- **Muon** (original) uses a normalized-momentum update direction from the SVD of the momentum buffer, enforcing search directions along the top singular vectors.
- **MuonH** solves a Hessian-free trust-region subproblem, with updates involving the rank decomposition and nuclear-norm scaling, standardizing step size and directionality [2502.02900].
- **AdamW** performs elementwise adaptive scaling via second-order moments, but does not globally control the layer capacity or spectrum. AdamW can induce norm growth and spectral concentration, problematic in pathological regimes such as grokking and heavy-tailed class distributions [2504.16041].

| Optimizer | Update Form        | Spectral Constraint | Preconditioning      |
|-----------|--------------------|--------------------|---------------------|
| Muon      | $-\eta U V^\top$   | $\|O\|_{2\to2}=1$  | Momentum / SVD      |
| MuonH     | $-\eta \|B\|_* U V^\top$ | Operator norm | Hessian-free, SVD   |
| AdamW     | $-\eta \, m/\sqrt{v+\epsilon}$ | None         | Diagonal, no SVD    |

## 5. Empirical Behavior: Stability under Heavy-Tail and Grokking

MuonH and the Muon family demonstrate practical advantages in domains with:
- **Heavy-tailed noise**: Empirically observed in large-scale training on non-uniform data distributions; MuonH's norm-based step mitigates gradient explosion and learning imbalance in rare/“tail” classes [2509.26030].
- **Grokking regime**: On modular arithmetic and parity tasks, Muon achieves a $\sim$33% reduction in mean grokking epoch (mem–gen transition) vs AdamW (102.89 vs 153.09), statistically highly significant ($t=5.0175$, $p=6.33\times10^{-8}$) [2504.16041].
- **Associative memory and transformers**: Experiments and theory show Muon’s update yields an isotropic singular value spectrum in critical associative-memory blocks (Value/Output attention, FFNs), ensuring balanced learning even for tail classes where Adam yields high disparity [2509.26030].

## 6. Hyperparameter Regimes and Implementation Notes

Robust operation of MuonH requires:
- *Hölder exponent* $\nu$: Typically $\nu=1$ (smooth ERM), but any $\nu\in(0,1]$ suffices.
- *Tail index* $p$: Empirically estimated or set to $2$ for bounded variance; $p\in(1,2]$ otherwise.
- *Step-size schedule* $\eta_t = \eta_0/(t+1)^a$, with $a > 1/(1+\nu)$. Practical $a \in [0.5,0.9]$.
- *Batch size*: Moderate constants (256–1024) or slow exponential increase.
- *SVD approximation*: 5 Newton–Schulz steps typically suffice, incurring negligible computational overhead relative to backpropagation.
- *Momentum*: $\beta \in [0.9,0.99]$. Additional terms in the convergence condition remain summable if $\sum_t \eta_t \beta^t < \infty$.

Recommended settings in neural language models: $\eta \approx 10^{-3}$, $\beta \approx 0.99$, spectral-norm bound $\sigma = 1.0$, and no weight decay on attention/FFN blocks for isolating effects [2603.15059][2504.16041][2509.26030].

## 7. Practical Significance and Research Directions

MuonH and related Muon optimizers provide algorithmic infrastructure for learning dynamics in nonconvex, nonsmooth, and statistically imbalanced regimes. Key advantages are:
- **Provably faster convergence in gradient norm** under minimal smoothness and with heavy-tailed stochastic effects.
- **Isotropic singular-value evolution** in critical network blocks, translating to improved learning for rare/“tail” data—a major advantage for long-tailed NLP and vision benchmarks.
- **Empirical acceleration** of delayed generalization transitions (grokking) and balanced performance across head and tail classes.
- **Layerwise normalization** preventing operator-norm blowup and aligning with implicit regularization trends seen empirically.

Ongoing research aims to integrate MuonH more closely with LLM pretraining pipelines, optimize SVD approximations further, and generalize convergence analysis to settings with additional nonlinear (e.g., batchnorm, attention) or structured noise [2603.15059][2502.02900][2509.26030].

Source: https://www.emergentmind.com/topics/muonh-optimizer