---
title: 'LoRA-Muon: Spectral Low-Rank Optimizer'
url: https://www.emergentmind.com/topics/lora-muon
type: topic
---

# LoRA-Muon: Spectral Low-Rank Optimizer

LoRA-Muon is an optimizer and fine-tuning paradigm that integrates the Muon optimizer—a spectral steepest-descent method—with the Low-Rank Adaptation (LoRA) framework. Originally motivated by the need for efficient and robust low-rank optimization in deep learning, LoRA-Muon offers unique advantages in stability, rank-invariant hyperparameter transferability, and spectral regularization. It addresses both the theoretical underpinnings of spectral descent on the low-rank manifold and practical limitations of existing factorwise adaptive optimizers. LoRA-Muon is especially salient for transfer and fine-tuning tasks, notably when navigating “optimizer mismatch” between Adam-pretrained models and spectral optimizers [2606.12921][2605.10468][2602.06385].

## 1. Mathematical Foundations: Spectral Steepest Descent on the Low-Rank Manifold

LoRA-Muon mathematically derives from steepest descent under the spectral norm constraint. For a loss $f:\mathbb{R}^{m\times n}\to\mathbb{R}$ with parameter matrix $W_t\in\mathbb{R}^{m\times n}$ and gradient $G_t=\nabla f(W_t)$, the spectral step is
\[
\Delta W_t^{\star} = -\eta\,\mathrm{msign}(G_t),
\]
where $\mathrm{msign}(G_t) = U V^T$ for $G_t = U\Sigma V^T$. In LoRA, the weight update is parameterized as $W = A B^T$, with $A\in\mathbb{R}^{m\times r}$, $B\in\mathbb{R}^{n\times r}$. The tangent space for updates is $T_W \mathcal{M}_r = \{\Delta A\,B^T + A\,\Delta B^T\}$, and LoRA-Muon splits the update between factors. Specializing to the spectral norm, the factorwise updates are
\begin{align*}
\Delta A &= -\frac{\eta}{2}\; \mathrm{msign}\left(\nabla_A f\,S_B^{-1/2}\right)\; S_B^{-1/2}\\
\Delta B &= -\frac{\eta}{2}\; \mathrm{msign}\left(\nabla_B f\,S_A^{-1/2}\right)\; S_A^{-1/2}
\end{align*}
with $S_A = A^T A$, $S_B = B^T B$, and $\mathrm{msign}$ implemented via Newton–Schulz polynomial iterations [2606.12921].

Compared to standard LoRA-Adam (which optimizes $A$ and $B$ with AdamW), this spectral structure leads to gauge invariance (updates depend only on $A B^T$, not the specific factorization) and a uniform spectral geometry aligned with those of full-rank Muon and Shampoo.

## 2. Comparison to Adam, Muon, and Optimizer Mismatch

Adam and Muon have fundamentally distinct implicit biases. Adam leverages per-parameter second moments and converges to minimum max-norm solutions, embedding a “max-norm” bias. Muon, via spectral normalization, equalizes step sizes across singular directions and induces a spectral-norm bias in the learned weights [2605.10468]. This dichotomy leads to optimizer mismatch: attempting Muon fine-tuning on Adam-pretrained parameters can degrade performance, as the spectral structure imposed by Muon disrupts the max-norm “shape” established by Adam.

Empirically, full fine-tuning with Muon on Adam-pretrained models results in a $+0.02$–$+0.03$ relative perplexity regression compared to Adam-matched runs, and a shifted and degraded learning-rate curve. This mismatch is proportional to update magnitude: stronger updates more readily override pretrained knowledge, exacerbating forgetting [2605.10468].

Low-rank adaptation with LoRA mitigates this mismatch, as the update is constrained to an $r$-dimensional subspace, inherently limiting the transformer’s deviation from its pretrained state. The worst-case inflation of mismatch is $\leq \sqrt{r}$ in the spectral norm, collapsing to no mismatch at $r=1$ and recovering full fine-tuning as $r \to n$.

## 3. Dynamical Properties: Uniform Spectral Growth and Global Convergence

A distinctive empirical and theoretical signature of LoRA-Muon is uniform spectral growth under spectral orthogonalization. In LLM fine-tuning, the singular values of $A B^T$ (i.e., the LoRA adapters) grow in parallel ("lock-step growth"), in stark contrast to AdamW, where larger modes dominate initially (“largest-first”). This effect, observed across practical settings (e.g., RoBERTa-Base and LLaMA-3.2-1B), persists with momentum and multi-factor extensions [2602.06385].

Theoretical analysis of the continuous-time spectral gradient flow (SpecGF) for LoRA factorization proves that all singular values $\sigma_i$ satisfy
\[
\dot{\sigma}_i(t) = 1 + o(1), \quad i = 1, \ldots, r
\]
for almost all initializations. With $\ell_2$ regularization, global convergence to the best rank-$r$ solution occurs almost surely. This spectral equalization is not present in AdamW or standard gradient flow, where rank learning proceeds unevenly [2602.06385].

## 4. Implementation: Algorithmic Formulation and Complexity

The LoRA-Muon algorithm avoids QR decompositions and the storage of second moments. The update steps consist of gradient evaluations, first-moment momentum, Gram matrix construction, inverse square-root computation (Newton–Schulz), and msign evaluations, all expressed as standard GEMMs plus $r \times r$ operations. Weight decay employs a split rule to match the desired linear dynamics:
\[
s = \sqrt{1 - \lambda\eta},\quad
A_{t+1} = s A_t + s^{-1} \Delta A_t,\quad
B_{t+1} = s B_t + s^{-1} \Delta B_t
\]
This split ensures that the low-rank update for $W$ aligns with the correct regularization schedule and preserves gauge symmetry [2606.12921].

Per-update computational cost and persistent state scale as $(m+n)r$ for LoRA-Muon versus $mn$ for dense Muon ($r \ll n$). Compared to other spectral low-rank methods (Spectron, LoRA-RITE), LoRA-Muon is both memory- and compute-efficient, and is invariant under arbitrary factor scaling, unlike Spectron [2606.12921].

### Table: Per-Update Characteristics for Various Methods

| Method           | Optimizer FLOPs           | Persistent State  |
|------------------|--------------------------|-------------------|
| Dense Muon       | $T_o(4nm^2+2m^3)$         | $mn$              |
| LoRA-Muon        | $(6+4T_o)(m+n)r^2\!+\!(16T_r+4T_o)r^3$ | $(m+n)r$      |
| Spectron         | $4T_o(m\!+\!n)r^2+\dots$  | $(m+n)r$          |

$T_o$ is the Newton–Schulz step count for msign; $T_r$ for inverse root.

## 5. Empirical Performance and Hyperparameter Recommendations

In compute-matched studies (e.g., TinyShakespeare), a rank-2 LoRA-Muon proxy attains the same optimal learning rate and validation loss as dense Muon, and rank-32 LoRA-Muon outperforms the dense baseline [2606.12921]. On extensive benchmarks spanning NLU (GLUE/T5-Base), NLG (Llama 2-7B), and vision (CLIP ViT-B/32), LoRA-Muon reliably matches or outperforms LoRA-Adam, closes the Adam–Muon gap, and halves optimizer state memory [2605.10468].

Learning-rate robustness is a key property: LoRA-Muon selects the same optimal linear adapter rate ($0.1$) across $r\in\{2,4,8,16,32\}$, width, depth, and gauge scaling. This is not the case for Adam-based factorwise optimizers, whose optimal rates degrade or fail to transfer. Moderate rank ($r \approx 8$–32) is recommended to balance expressiveness with mismatch suppression. For regularization, split weight decay with $\lambda=0.01$ matched to dense Muon is effective.

## 6. LoRA-Muon in Fine-Tuning and Optimizer Mismatch Mitigation

When fine-tuning Adam-pretrained models, naive application of Muon induces optimizer mismatch and catastrophic forgetting. LoRA-Muon constrains update magnitude, preserves Adam-learned structure, and thus sharply reduces the mismatch both theoretically and empirically. On tasks where the mismatch is severe (e.g., MetaMath $\rightarrow$ GSM8K), LoRA-Muon outperforms LoRA-Adam at low and moderate ranks; where mismatch is mild, LoRA-Muon matches or slightly exceeds Adam across all ranks. LoRA-Muon also incurs less catastrophic forgetting than full Muon fine-tuning.

Variants of LoRA (rsLoRA, LoRA-One, PiSSA, AdaLoRA, LoRA-RITE, DoRA) improve LoRA-Adam to a degree, but optimizer-agnostic enhancements do not further boost LoRA-Muon’s peak accuracy. Techniques that amplify update norms can reinstate mismatch effects.

**Practical guidance:** Always use LoRA when fine-tuning Adam-pretrained models with Muon. Tune learning rates empirically, as optimal rates do not transfer from Adam to Muon. Recalibrate rank and hyperparameters; avoid transplanting LoRA variants tuned for Adam without verification on Muon [2605.10468].

## 7. Limitations and Future Directions

Current empirical validation is primarily at the TinyShakespeare and moderate-scale NLU/NLG/vision benchmark levels. Full LLM pretraining, downstream fine-tuning, and large-scale distributed settings remain to be systematically tested. The transfer of full hyperparameter landscapes beyond learning rate (e.g., joint $\lambda$–$\eta$ sweeps) and exploration of further unitary-invariant norm adaptations (e.g., Ky-Fan) or adaptive second-order moment transport on the low-rank manifold are open areas. Interactions with quantization regimes, mixed precision, and distributed computation are unexplored [2606.12921].

A plausible implication is that the spectral geometry and optimizer-invariant properties of LoRA-Muon could yield increased resilience to initialization pathologies, scaling instabilities, and irregular optimization landscapes that hamper existing LoRA-Adam implementations, particularly as model size and heterogeneity grow.

---

**References**  
- "LoRA-Muon: Spectral Steepest Descent on the Low-Rank Manifold" [2606.12921]  
- "Can Muon Fine-tune Adam-Pretrained Models?" [2605.10468]  
- "Uniform Spectral Growth and Convergence of Muon in LoRA-Style Matrix Factorization" [2602.06385]

Source: https://www.emergentmind.com/topics/lora-muon