Papers
Topics
Authors
Recent
Search
2000 character limit reached

The Loss Does Not See the Basis, but Adam Does

Published 5 Aug 2026 in cs.LG, math.OC, and stat.ML | (2608.05136v1)

Abstract: Gradient descent on a factored model W=UV<sup>W = UV<sup>\top is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization, is not. We trace the difference to the gauge symmetry of the loss, its invariance under (U,V)(UQ,VQ)(U, V) \mapsto (UQ, VQ). Gradient flow's low-rank mechanism is available to an optimizer only if that optimizer is gauge-equivariant, a condition necessary for the transfer but not sufficient for low-rank recovery. Gradient descent, momentum, "shared-scalar" Adam, Muon, and Shampoo satisfy it. Adam, RMSProp, and the other coordinate-wise methods do not. A structure theorem characterizes the memoryless equivariant rules as exactly the Gram-determined left preconditioners, and a transfer theorem carries gradient flow's pathwise properties to common-scalar flows. We then sort nine update rules on underdetermined matrix sensing by recovery error against the planted ground truth. A one-parameter family from coordinate-wise to shared-scalar preconditioning restores the bias monotonically, isolating anisotropy as the cause. A "spectral schedule" reconciles two opposing reports about Muon: equal-rate updates recover exactly low-rank targets but lose their edge as the spectral tail grows. In transformers, Adam separates two gauge-equivalent initializations at the first step, where the equivariant optimizers stay at float precision, and ends with the per-head invariants WQ<sup></sup>WKW_Q<sup>\top</sup> W_K 56% apart in relative Frobenius distance, a gap no per-head rotation can close. On two hyperspectral datasets at matched training loss, gradient descent cuts held-out error by 43-44% at the lowest sampling density, and at lower effective rank. Basis choice is therefore not a tuning detail but a decision about which interpolant the optimizer selects.

Authors (1)

Summary

  • The paper demonstrates that Adam’s coordinate-wise preconditioning breaks gauge equivariance, producing higher-rank interpolants and substantially poorer recovery than GD, scalar-Adam, Muon, and Shampoo in matrix-sensing experiments.
  • The paper separates symmetry compatibility from spectral scheduling, showing that equivariance preserves access to gradient-flow-like bias while mode-growth dynamics determine whether an optimizer favors exactly low-rank or spectrally distributed targets.
  • The paper finds that reducing Adam’s preconditioner anisotropy improves recovery and rank, while gauge-aware updates also reduce basis-dependent divergence in attention models and improve generalization on hyperspectral completion benchmarks.

The Loss Does Not See the Basis, but Adam Does

Central thesis

“The Loss Does Not See the Basis, but Adam Does” (2608.05136) studies how optimizer geometry affects implicit bias in overparameterized factored models. Its central claim is that the familiar low-rank preference of gradient-based optimization is not determined solely by the loss landscape or by interpolation. It also depends on whether the optimizer respects the internal gauge symmetry of the factorization.

For a factorized matrix W=UVW = UV^\top, the transformation

(U,V)(UQ,VQ),QO(k),(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),

leaves WW and therefore any loss of the form L(U,V)=f(UV)L(U,V)=f(UV^\top) unchanged. The matrices UU and VV can consequently be represented in infinitely many orthogonal latent bases without changing the represented function. The paper calls this transformation the gauge symmetry. The optimizer, however, acts on UU and VV separately rather than directly on WW. The decisive question is therefore whether its update commutes with this gauge action.

The paper’s principal conclusion is that coordinate-wise adaptive methods, especially Adam, break this symmetry, whereas GD, momentum, scalar-preconditioned Adam, Muon, and Shampoo preserve it under the stated conditions. This distinction predicts which interpolating solution is selected at finite computational budgets. In the experiments, symmetry-preserving methods retain substantially more of the low-rank inductive bias associated with gradient flow, while coordinate-wise methods generally converge to higher-rank interpolants with poorer recovery.

Gauge equivariance as the organizing principle

The paper formalizes gauge equivariance for optimizers with internal state. If an initial factorization is transformed by QQ, an equivariant optimizer must produce a correspondingly transformed trajectory, including an appropriate transformation of its state. Consequently, the represented matrices (U,V)(UQ,VQ),QO(k),(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),0 remain identical for all gauge-equivalent initializations.

The distinction is immediate for the first Adam update. With zero moment estimates and bias correction, Adam applies the entrywise map

(U,V)(UQ,VQ),QO(k),(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),1

In general,

(U,V)(UQ,VQ),QO(k),(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),2

Thus, Adam can produce different represented matrices after a single optimization step even when two initial parameterizations implement exactly the same function. This is stronger than a difference in optimizer state or in the factor representation: the function represented by the two trajectories itself diverges.

The paper proves a broader classification result for memoryless updates. An update (U,V)(UQ,VQ),QO(k),(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),3 is gauge-equivariant if and only if it can be written as

(U,V)(UQ,VQ),QO(k),(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),4

where (U,V)(UQ,VQ),QO(k),(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),5 depends only on the gauge-invariant Gram matrix (U,V)(UQ,VQ),QO(k),(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),6. This theorem excludes fixed nonlinear coordinate-wise maps except for linear rescalings. It also explains why matrix-structured methods such as Muon and Shampoo can preserve the symmetry despite violating properties often associated with gradient flow, such as balancedness.

Figure 1

Figure 1: Gauge equivariance separates optimizers according to whether their updates preserve the available symmetry of the factorized representation.

The paper carefully distinguishes equivariance from low-rank recovery. Equivariance is necessary for directly transferring gradient-flow behavior across gauge-equivalent parameterizations, but it is not sufficient to guarantee a low-rank solution. The remaining degree of freedom is the optimizer’s spectral schedule: the relative rates at which different singular modes grow. ScaledGD is given as an equivariant counterexample that equalizes mode growth and therefore does not reproduce the greedy low-rank dynamics of gradient flow. Conversely, a long-annealed sign-based method can recover a low-rank target despite breaking equivariance. The correct conceptual decomposition is therefore two-dimensional: gauge compatibility determines access to the symmetry-consistent bias, while spectral dynamics determine the specific bias within the equivariant class.

Transfer from gradient flow

A transfer theorem establishes that continuous-time dynamics with a positive shared scalar preconditioner follow exactly the same path as gradient flow under a reparameterization of time. If

(U,V)(UQ,VQ),QO(k),(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),7

with (U,V)(UQ,VQ),QO(k),(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),8 determined from gauge-invariant trajectory statistics, then the change of variables

(U,V)(UQ,VQ),QO(k),(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),9

produces ordinary gradient flow in the variable WW0. Therefore, pathwise and limit-point results established for gradient flow can be transferred to common-scalar preconditioned flows, subject to the required divergence and convergence assumptions.

This result does not formally cover scalar-Adam with its first-moment EMA, nor does it cover arbitrary equivariant stateful optimizers. The paper consequently treats the close empirical agreement between scalar-Adam and GD as evidence rather than as a complete theorem. It also emphasizes that rate-dependent claims, such as convergence speed, do not transfer under time reparameterization.

Matrix-sensing experiments

The primary controlled experiment uses underdetermined Gaussian matrix sensing. The planted matrix is WW1 with rank WW2, the factors are overparameterized with latent dimension WW3, and the number of measurements is twice the rank-WW4 degrees of freedom. All methods start from small initialization, use no weight decay, and are trained until the measurement residual falls below WW5. This protocol removes residual training error as an explanation for differences in recovery.

The results produce a pronounced separation. The equivariant methods obtain recovery errors between WW6 and WW7, whereas the coordinate-wise methods obtain errors between WW8 and WW9. Muon achieves near-exact recovery with error L(U,V)=f(UV)L(U,V)=f(UV^\top)0 and effective rank L(U,V)=f(UV)L(U,V)=f(UV^\top)1. GD obtains recovery error L(U,V)=f(UV)L(U,V)=f(UV^\top)2, scalar-Adam L(U,V)=f(UV)L(U,V)=f(UV^\top)3, and Shampoo L(U,V)=f(UV)L(U,V)=f(UV^\top)4. In contrast, Lion, signum, RMSProp, Adafactor, and Adam obtain errors of approximately L(U,V)=f(UV)L(U,V)=f(UV^\top)5, L(U,V)=f(UV)L(U,V)=f(UV^\top)6, L(U,V)=f(UV)L(U,V)=f(UV^\top)7, L(U,V)=f(UV)L(U,V)=f(UV^\top)8, and L(U,V)=f(UV)L(U,V)=f(UV^\top)9, respectively.

Figure 2

Figure 2: Recovery performance on the underdetermined sensing problem separates gauge-equivariant optimizers from coordinate-wise adaptive methods at matched interpolation.

Because every method interpolates, the performance gap reflects interpolant selection rather than optimization accuracy on the observed measurements. The result also exposes a limitation of interpreting effective rank in isolation. Signum produces a lower effective rank than several other coordinate-wise methods but still has poor recovery, indicating that it can concentrate updates in incorrect directions. Low rank in parameter space is not equivalent to recovery of the planted low-rank matrix.

The convex minimum-nuclear-norm solution obtains recovery error UU0, outperforming the practical GD trajectory at the reported finite step size. This difference is consistent with the restricted nature of the gradient-flow theory: infinitesimal initialization, continuous time, and asymptotic convergence are not identical to finite-step GD at initialization scale UU1. Muon nevertheless achieves better recovery than the nuclear-norm reference on the reported task, demonstrating that its equal-rate spectral dynamics encode a bias that is not reducible to nuclear-norm minimization.

The ranking persists in an H100 replication ladder from UU2 through UU3 using ten seeds. GD, scalar-Adam, and Muon consistently outperform the clean coordinate-wise methods, with Muon remaining essentially exact through the tested sizes. The authors appropriately qualify the result: Shampoo’s fixed damping becomes poorly calibrated at larger dimensions, and prolonged annealing can enable signum to recover despite its non-equivariance. Thus, the two-cluster structure is a finite-budget empirical classification, not an asymptotic theorem covering every hyperparameter regime.

The preconditioner-anisotropy dial

To isolate the mechanism within Adam, the paper introduces Adam-UU4, a continuous family interpolating between standard Adam and scalar-Adam. At UU5, the denominator is coordinate-wise; at UU6, it is a shared gauge-invariant RMS scalar. Momentum, second-moment estimation, learning-rate grids, and other algorithmic components remain fixed while only the anisotropy of the denominator changes.

The recovery and effective rank improve monotonically as UU7 decreases. In the envelope experiment, recovery changes from UU8 at UU9 to VV0, VV1, VV2, and VV3 at VV4, VV5, VV6, and VV7, respectively. Effective rank simultaneously decreases from VV8 to $5.4.

Figure 3

Figure 3: Reducing Adam’s coordinate-wise anisotropy continuously restores lower-rank recovery while preserving the remaining adaptive and momentum mechanisms.

This experiment is particularly important because it rules out several alternative explanations. The restoration is not caused merely by removing adaptivity, eliminating momentum, changing the global learning-rate scale, or altering the training loss. The only systematic change is the degree to which the preconditioner distinguishes coordinates in a basis-dependent manner. The result converts the binary symmetry argument into a quantitative mechanism: the stronger the anisotropy, the more strongly the optimizer reads the arbitrary latent basis.

The trade-off is computational. On the sensing task, moving from standard Adam to scalar-Adam increases the number of steps required for interpolation by approximately a factor of eight, from roughly $V$9 to $U$0. Adam’s coordinate-wise geometry therefore provides optimization advantages in exchange for changing the implicit selection rule.

Attention-head experiments

The same gauge structure appears inside attention heads. For query and key matrices, the logits depend on the product

$U$1

An orthogonal transformation applied jointly to the head representations,

$U$2

preserves this product and hence preserves the model function at initialization.

The paper trains gauge-equivalent transformer twins on modular addition. Under Adam, the twins separate at the first update: their relative logit distance reaches $U$3 after one step, whereas a same-basis perturbation of magnitude $U$4 produces only $U$5 of drift. The Adam twins ultimately exhibit a relative Frobenius discrepancy of $U$6 in the invariant $U$7, and their final relative logit distance reaches approximately $U8.</p><p><imgsrc="https://images.emergentmind.com/paperimages/260805136/phasediagram.png"alt="Figure4"title=""class="markdownimage"loading="lazy"></p><p><pclass="figurecaption">Figure4:Adamproducesanimmediatestructuralsplitbetweengaugeequivalentattentionmodels,whileequivariantmethodsremainatthenumericalnoisefloorduringtheinitialupdate.</p></p><p>Heavyball<ahref="https://www.emergentmind.com/topics/stochasticgradientdescentsgd"title=""rel="nofollow"dataturbo="false"class="assistantlink"xdataxtooltip.raw="">SGD</a>andscalarAdamremainindistinguishabletofloatingpointprecisionatthefirststep.Muonalsoremainsgaugeequivariantinexactarithmetic,butitspracticalNewtonSchulzapproximationisnumericallyunstablenearsmallsingularvalues;itsgaugeandnoisetwinsfollowsimilartrajectories,identifyingnumericalchaosratherthanstructuralbasisdependenceasthesourceofdivergence.</p><p>TheauthorsreproducethefirststepAdamsplitinlargerfloat64transformersandinastochasticcharacterlevelLLM.Inthelatter,Adamtwinsdivergebyapproximately8.</p> <p><img src="https://images.emergentmind.com/paper-images/2608-05136/phase_diagram.png" alt="Figure 4" title="" class="markdown-image" loading="lazy"></p> <p><p class="figure-caption">Figure 4: Adam produces an immediate structural split between gauge-equivalent attention models, while equivariant methods remain at the numerical-noise floor during the initial update.</p></p> <p>Heavy-ball <a href="https://www.emergentmind.com/topics/stochastic-gradient-descent-sgd" title="" rel="nofollow" data-turbo="false" class="assistant-link" x-data x-tooltip.raw="">SGD</a> and scalar-Adam remain indistinguishable to floating-point precision at the first step. Muon also remains gauge-equivariant in exact arithmetic, but its practical Newton–Schulz approximation is numerically unstable near small singular values; its gauge and noise twins follow similar trajectories, identifying numerical chaos rather than structural basis dependence as the source of divergence.</p> <p>The authors reproduce the first-step Adam split in larger float64 transformers and in a stochastic character-level LLM. In the latter, Adam twins diverge by approximately U$9–$V$0 after one step despite receiving identical minibatch streams, while equivariant methods remain at approximately $V$1–$V$2. The twins achieve similar validation losses but different functions. This result demonstrates that optimizer-induced basis dependence is not confined to matrix-sensing toy problems.

The attention experiments also have implications for model merging. Alignment procedures can remove differences related to internal rotations, but they cannot reconcile disagreement in the invariant $V$3. If Adam causes two gauge-equivalent runs to acquire different invariants, no per-head Procrustes alignment can make them functionally identical. This does not constitute an end-to-end model-merging experiment, but it identifies a structural obstruction to alignment-based merging.

Spectral schedules and the Muon boundary

The paper separates optimizer class from spectral schedule using targets with a controlled spectral tail. The target contains a rank-$V$4 component and an orthogonal dense tail carrying energy $V$5. Muon is nearly exact at $V$6, but its advantage disappears as the tail grows. The crossover occurs near $V$7, corresponding to approximately $V8tailenergy.</p><p><imgsrc="https://images.emergentmind.com/paperimages/260805136/realdatatrajectory.png"alt="Figure5"title=""class="markdownimage"loading="lazy"></p><p><pclass="figurecaption">Figure5:MuonsequalratespectralscheduleisoptimalforexactlylowranktargetsbutbecomeslessrobustthanGDwhenthetargetcontainsanontrivialspectraltail.</p></p><p>At8 tail energy.</p> <p><img src="https://images.emergentmind.com/paper-images/2608-05136/realdata_trajectory.png" alt="Figure 5" title="" class="markdown-image" loading="lazy"></p> <p><p class="figure-caption">Figure 5: Muon’s equal-rate spectral schedule is optimal for exactly low-rank targets but becomes less robust than GD when the target contains a nontrivial spectral tail.</p></p> <p>At V$9, recovery errors are approximately $W$0 for Muon, $W$1 for GD, $W$2 for Shampoo, and $W$3 for Adam. At $W$4, GD and Muon are statistically comparable, with errors near $W$5. At $W$6, GD becomes preferable, obtaining approximately $W$7 versus Muon’s $W$8. At $W$9, the problem approaches a measurement-determined floor and the ordering becomes less informative.

The proposed explanation is dynamical. GD preferentially amplifies large singular modes, producing a greedy schedule that suppresses weak modes during the time required to fit the dominant structure. Muon’s approximately equal-rate schedule grows all modes more uniformly. This is beneficial when every nonzero mode belongs to the true low-rank signal, but harmful when the target contains weak or poorly identified spectral components that should not be fitted aggressively.

A decoupled two-timescale analysis formalizes this contrast. Under greedy dynamics, the weaker mode remains bounded by a term proportional to

$Q$0

while the initialization scale tends to zero. Under equal-rate dynamics, the weaker mode is fitted at its own saturation value by the time the leading mode reaches a prescribed threshold. This analysis explains why Muon can simultaneously exhibit a strong simplicity bias on exactly low-rank data and poor behavior on mixed-spectrum targets.

Real-data validation

The paper evaluates the mechanism on Indian Pines and Pavia University hyperspectral matrix-completion benchmarks. The factorization rank is $Q$1, while the intrinsic spectral rank is estimated as $Q$2. Comparisons are made at matched training losses rather than matched numbers of iterations, avoiding the confound that a slower optimizer may appear better simply because it has not yet interpolated.

At sampling density $Q$3, GD reduces held-out RMSE by $Q$4 on Indian Pines and $Q$5 on Pavia University relative to Adam. At density approximately $Q$6, the reductions are $Q$7 and $Q$8, respectively. On Indian Pines at the lower density, GD obtains RMSE $Q$9 versus Adam’s $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$00, with effective rank near $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$01 versus Adam’s $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$02.

The trajectories show that Adam’s held-out error can increase as its training loss decreases. On Indian Pines, Adam’s error rises from approximately $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$03 at training loss $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$04 to $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$05 at $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$06, while effective rank increases from $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$07 to $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$08. GD continues to improve. This is direct evidence that, under underdetermination, deeper interpolation with a basis-dependent optimizer can worsen generalization.

Muon does not uniformly improve real-data performance. On Indian Pines it performs worse than Adam during most of the matched-loss trajectory and retains effective rank close to the model capacity until late training. On Pavia University it eventually becomes competitive or superior to Adam after a late rank collapse. These results reinforce the two-axis interpretation: gauge equivariance makes a low-rank mechanism available, but the spectral schedule determines whether that mechanism is suitable for the data.

The size of the GD–Adam gap decreases with sampling density, approaching zero as the observations increasingly constrain the matrix. This behavior is theoretically expected: when the interpolation set becomes narrower, optimizer-dependent interpolant selection has less room to affect the final solution.

Repairing FlowAdam

The paper applies its criterion retrospectively to FlowAdam, an optimizer that injects a clipped gradient-flow velocity into Adam. The original coordinate-wise clipping operation itself breaks gauge equivariance. Replacing coordinate-wise clipping with global-norm clipping improves recovery from $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$09 to $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$10, a $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$11 reduction in the reported experiment.

Combining global-norm clipping with the Adam-$(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$12 dial yields recovery error $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$13, compared with $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$14 for the dial alone and $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$15 for Adam. The improvement is approximately $(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),$16 at the fixed learning rate used in the comparison. However, the paper reports that the advantage largely disappears on finer learning-rate grids, indicating a finite-step acceleration effect rather than a distinct asymptotic implicit bias.

This case study illustrates the practical value of the symmetry criterion. It can identify a hidden failure mode in an optimizer whose intended purpose is to restore gradient-flow-like behavior. It also suggests a general design rule: clipping or normalization should act on invariant global or matrix-structured quantities rather than independently on coordinates whenever the parameterization contains a nontrivial gauge.

Theoretical and practical implications

The theoretical contribution is a refinement of the standard account of implicit regularization in factored models. The paper argues that low-rank bias should not be described only in terms of initialization, factorization depth, or the loss. The optimizer’s transformation properties are part of the inductive bias. Coordinate-wise adaptive methods impose an extrinsic geometry that is not intrinsic to the represented matrix and can therefore alter the selected interpolant.

The structure theorem offers a constructive framework for designing symmetry-compatible optimizers. A permissible memoryless update can depend on the gradient through gauge-invariant Gram statistics and then apply a left preconditioner. Stateful extensions can use covariant momentum, invariant left accumulators, and conjugation-equivariant right accumulators. This perspective encompasses GD, scalar-Adam, Muon, Shampoo, and related matrix optimizers within a common formalism.

The results do not imply that Adam should be replaced universally. Coordinate-wise adaptivity can be valuable in stiff, heterogeneous, heavy-tailed, or poorly conditioned optimization problems. The paper explicitly identifies settings in which anisotropic methods may be preferable, including PINNs and tasks requiring stronger or more targeted regularization. The relevant practical question is not whether an optimizer is equivariant in the abstract, but whether the desired inductive bias is compatible with the problem’s spectral structure and data constraints.

Several limitations remain. The transfer theorem is restricted to common-scalar continuous-time dynamics and does not yet characterize momentum-augmented scalar-Adam. The interaction between gauge symmetry and minibatch noise is only empirically examined. The Muon phase boundary lacks a complete coupled-dynamics theorem. Finally, the long-annealed signum exception shows that non-equivariance is not a logically necessary barrier to low-rank recovery; it is a mechanism-level predictor whose scope depends on the optimization regime.

Future developments are likely to combine the two axes identified by the paper: explicit control of symmetry compatibility and explicit control of spectral transfer functions. Fractional spectral methods, equivariant adaptive optimizers, and gauge-aware preconditioners could interpolate between greedy GD-like dynamics and equal-rate polar dynamics. In attention architectures, symmetry-aware optimization may also become relevant for reproducible fine-tuning, model merging, and parameter-space comparisons. A complete theory will need to treat stateful stochastic dynamics on quotient spaces, finite-precision instability, and interactions between architectural symmetries and optimizer-induced geometries.

Conclusion

“The Loss Does Not See the Basis, but Adam Does” (2608.05136) presents gauge equivariance as a precise criterion for analyzing optimizer-dependent interpolant selection in factored models. Adam’s coordinate-wise preconditioner breaks the orthogonal gauge symmetry at the first update, while GD, scalar-Adam, Muon, and Shampoo preserve it under the paper’s formal conditions. On underdetermined sensing and hyperspectral completion, this distinction is associated with markedly different rank profiles and recovery errors, including (U,V)(UQ,VQ),QO(k),(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),17–(U,V)(UQ,VQ),QO(k),(U,V) \mapsto (UQ,VQ), \qquad Q \in O(k),18 held-out improvements for GD over Adam at the lowest sampling densities.

The paper’s broader contribution is to separate symmetry preservation from spectral scheduling. Equivariance determines whether an optimizer respects the intrinsic representation geometry; spectral dynamics determine which structures it amplifies. This distinction explains both the failure of Adam to preserve gradient-flow low-rank bias and the opposing behavior of Muon on exactly low-rank versus spectrally tailed targets. It also provides concrete design principles for constructing optimizers whose implicit bias is tied to the represented function rather than to an arbitrary parameter basis.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.