Papers
Topics
Authors
Recent
Search
2000 character limit reached

Complete-muE: Optimal Hyperparameter Transfer and Scaling for MoE Models

Published 22 May 2026 in cs.LG | (2605.23893v1)

Abstract: We propose Complete-muE, a framework which targets hyperparameter transfer across dense FFN and any Mixture-of-Experts (MoE) setups in transformer blocks. Existing tools such as μμP (requires fixed architectue) or SDE (requires fixed per-step token count) cannot directly solve the hyperparameter transfer problem in MoE setups because Dense to MoE transfer or MoE total experts scaling changes both architecture and tokens per expert. Complete-muE solves this challenge with a two-bridge system: Bridge~I maps between dense FFN and Dense MoE by active-width μμP with a normalized router scale. Bridge~II maps between Dense MoE and sparse MoE by activated-expert scaling, where the first-order SDE LR/WD correction cancels while a bounded residual σ0σ_0 shift remains. The resulting transfer rule, which we term as Complete muE, covers changes in activated experts, total capacity, granularity, and shared/group-balanced hybrids for MoE models as well as network width/depth, batch size, and duration changes for general Transformer models. Extensive LLM and diffusion model pretraining experiments confirm that complete-muE yields relatively stable hyperparameter optima across model architectures and parameter counts -- with only minor drift consistent with the non-strict SDE behavior of Bridge~II. In practice this drift is small enough that hyperparameters tuned on a single dense reference transfer near-optimally to all MoE configurations -- \emph{tune dense once, transfer to all} is the practical recipe at the core of Complete-muE. This enables MoE models to achieve accelerated convergence speedup over dense models when scaling model capacity without costly hyperparameter search.

Summary

  • The paper introduces two composable transfer bridges based on active width and expert workload, allowing dense-tuned AdamW settings to transfer across activated experts, capacity, granularity, depth, and routing variants.
  • Experiments show broad alignment of low-loss regions and consistent gains without per-architecture retuning, including up to 5.5× faster convergence than dense models and improved downstream performance from 44.3 to 50.6 on a 13-task average.
  • The method remains approximate because expert-side signal-to-noise shifts, routing imbalance, finite-width effects, and optimizer assumptions can cause hyperparameter drift, while capacity scaling generally offers lower latency than granularity scaling.

Motivation and problem statement

Hyperparameter transfer for Mixture-of-Experts (MoE) transformers has been fragmented across two incompatible toolchains. μ\muP-style parameterizations transfer initialization and optimizer settings across model-size changes but assume a fixed architecture, while SDE-based rules transfer across token batch size and training duration for a fixed parameterized model. Sparse MoE layers violate both assumptions simultaneously: routing changes the architecture relative to a dense FFN, and it also changes the per-expert token workload, since under balanced routing each expert processes roughly Ba/NBa/N tokens per step (BB is the global batch, NN the total expert count). Dense-to-sparse transfer and total-expert scaling therefore couple architecture transfer with workload transfer, and neither existing framework covers the composition.

Complete-muE addresses this gap by decomposing any FFN/MoE transfer into two primitive "bridges" composed through a single governing quantity, the active width Ha=ahH_a = ah (activated experts aa, per-expert width hh). The practical payoff claimed by the authors is strong: hyperparameters tuned once on a dense reference FFN transfer near-optimally to every MoE configuration—"tune dense once, transfer to all."

The two-bridge construction

The paper writes all FFN/MoE variants in one form, y(x)=A(Hact)∑igi(x) oi(x)y(x) = A(H_{\text{act}})\sum_i g_i(x)\, o_i(x), where gi(x)g_i(x) specializes to an always-on block (dense FFN), all-experts-active normalized routing (Dense MoE), or top-aa normalized routing (sparse MoE). Routing is restricted to token-choice (top-Ba/NBa/N0), which fixes the active set size deterministically per token; this makes the Ba/NBa/N1P update-size matching exact and avoids the train–inference mismatch of expert-choice routing.

Bridge I (dense FFN ↔ Dense MoE) applies active-width Ba/NBa/N2P at fixed backbone width Ba/NBa/N3: output multiplier Ba/NBa/N4, down-projection init std Ba/NBa/N5, and unchanged down-projection learning rate. Normalized routing would otherwise shrink the routed update by Ba/NBa/N6, so Complete-muE introduces a route scale Ba/NBa/N7. An optional finite-width correction Ba/NBa/N8 exists but is set to 1, justified by an expansion showing Ba/NBa/N9 stays order-one near uniform logits.

Bridge II (Dense MoE ↔ sparse MoE) handles changes in activated experts. After layer-level matching via BB0, the remaining effect is stochastic: each expert's batch and duration both scale as BB1. The first-order SDE LR/WD multiplier BB2 therefore cancels exactly, leaving BB3, BB4. The paper is explicit that this is not a strict SDE invariance: the signal-to-noise parameter shifts as BB5, which is not absorbed by the first-order correction, so mild bounded hyperparameter drift across BB6 is expected theoretically and observed empirically. Larger BB7 improves expert-side SNR and lowers attainable loss even at unchanged raw hyperparameters—an important consequence, since it means sparsity reduction yields quality gains without retuning.

All remaining cases are compositions rather than new primitives. Capacity scaling (BB8 at fixed BB9) decomposes into a Dense-MoE width step followed by reverse sparsification; the width factors cancel, recovering the same rule up to NN0, again with mild drift from the non-strict Bridge II behavior. Granularity scaling at fixed density NN1 reduces exactly to Dense-MoE width transfer because expert batch and duration (NN2, NN3) are invariant to the partition NN4. Shared, group-balanced, and hybrid blocks are treated as one expanded FFN of total active width NN5 with a single global route scale on routed groups only. Global AdamW factors compose separately: schedule changes (batch/duration) apply the standard NN6 rule uniformly—exact when total tokens are fixed, approximate with a NN7 shift when iterations are fixed—while MoE architectural changes require no additional global multiplier.

Empirical validation

The evaluation uses controlled proxy sweeps (LM: NN8, 32 layers, 25k steps; diffusion transformer: flow matching, 100k steps) plus large-scale runs. Across activated-expert counts (64e2a–64e16a), capacity, granularity, shared experts, group-balanced routing, depth, width, and FFN expansion ratio, the low-loss regions for LR, weight decay, and init std remain broad and aligned, with representative stable ranges such as LR NN9–Ha=ahH_a = ah0 and WD Ha=ahH_a = ah1–Ha=ahH_a = ah2 for LM. The most direct evidence for the recipe fixes AdamW hyperparameters at dense-tuned values and varies only the MoE architecture along four axes: loss decreases consistently for both LM and diffusion without per-setting retuning. Notably, increasing shared experts consistently raises loss, leading the authors to conclude that shared experts are unnecessary for quality unless needed for compute–communication overlap—a claim that runs against the prevalence of shared experts in recent production systems.

A single-H100 latency benchmark quantifies the scaling trade-off between capacity and granularity. Capacity scaling from 8 to 256 experts costs only Ha=ahH_a = ah3–Ha=ahH_a = ah4 over the dense baseline (87.8–97.0 ms vs. 81.0 ms/step), whereas granularity scaling raises latency from 84.1 to 135.0 ms/step, and dense width scaling reaches 667.6 ms/step before exhausting memory. Capacity is therefore the cheaper axis at large per-device token batches.

Large-scale runs apply one hyperparameter setting per modality family—four diffusion regimes (256P/512P images, 240P key frames, 240P 5s videos) share LR Ha=ahH_a = ah5, WD Ha=ahH_a = ah6, and the LM run uses LR Ha=ahH_a = ah7, WD Ha=ahH_a = ah8—with backbone width Ha=ahH_a = ah9 the proxy width (aa0). At roughly 0.62B active / 6.29B total parameters and 100k steps, MoE achieves approximately aa1 convergence speedup on 256P images, aa2 on 240P 5s video, and aa3–aa4 on LLM training relative to dense during the stable-LR phase. Downstream benchmarks improve the 13-task average from 44.3 (dense) to 49.1 (128e8a4g1s) and 50.6 (128e8a1s); the non-grouped variant leads on average, consistent with smaller-scale findings that group-balanced routing slightly lags.

Limitations and open questions

The paper concedes several boundaries explicitly. First, Bridge II is not an exact invariance: the residual aa5 shift produces bounded but nonzero hyperparameter drift across activated experts, capacity, and fixed-iteration batch changes, so the "near-optimal" transfer claim rests on empirical smallness of this drift rather than theory. Second, the cancellation argument assumes approximate load balancing; the imbalanced-routing extension shows the cancellation survives only at the expert-averaged level, with load fluctuations contributing to drift. Third, the route-scale rule discards the finite-width factor aa6, which can deviate from one if selected router logits become highly separated during training—the analysis covers initialization but not fully trained routers. Fourth, the capacity–noise trade-off means increasing aa7 need not monotonically improve loss; saturation or reversal is possible once the SNR penalty aa8 dominates, and the paper does not characterize where this crossover occurs. Finally, the framework is restricted to token-choice routing and the AdamW gradient-magnitude-normalized regime; extension to expert-choice routing or other optimizers remains open.

Conclusion

Complete-muE provides a compositional AdamW transfer rule covering activated experts, total capacity, fixed-density granularity, shared/group-balanced hybrids, and standard width/depth/batch/duration changes, built from two bridges whose key insight is that expert-side batch and duration ratios cancel in the first-order SDE correction. Controlled sweeps and large-scale multimodal/LM runs support the central operational claim—that a single dense calibration suffices for near-optimal MoE training—with convergence speedups up to aa9 over dense baselines at matched active parameter count. The residual theoretical gap between the demonstrated "relatively stable" transfer and strict invariance, and its dependence on load-balancing assumptions, remain the main open questions.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.