- The paper introduces two composable transfer bridges based on active width and expert workload, allowing dense-tuned AdamW settings to transfer across activated experts, capacity, granularity, depth, and routing variants.
- Experiments show broad alignment of low-loss regions and consistent gains without per-architecture retuning, including up to 5.5× faster convergence than dense models and improved downstream performance from 44.3 to 50.6 on a 13-task average.
- The method remains approximate because expert-side signal-to-noise shifts, routing imbalance, finite-width effects, and optimizer assumptions can cause hyperparameter drift, while capacity scaling generally offers lower latency than granularity scaling.
Motivation and problem statement
Hyperparameter transfer for Mixture-of-Experts (MoE) transformers has been fragmented across two incompatible toolchains. μP-style parameterizations transfer initialization and optimizer settings across model-size changes but assume a fixed architecture, while SDE-based rules transfer across token batch size and training duration for a fixed parameterized model. Sparse MoE layers violate both assumptions simultaneously: routing changes the architecture relative to a dense FFN, and it also changes the per-expert token workload, since under balanced routing each expert processes roughly Ba/N tokens per step (B is the global batch, N the total expert count). Dense-to-sparse transfer and total-expert scaling therefore couple architecture transfer with workload transfer, and neither existing framework covers the composition.
Complete-muE addresses this gap by decomposing any FFN/MoE transfer into two primitive "bridges" composed through a single governing quantity, the active width Ha=ah (activated experts a, per-expert width h). The practical payoff claimed by the authors is strong: hyperparameters tuned once on a dense reference FFN transfer near-optimally to every MoE configuration—"tune dense once, transfer to all."
The two-bridge construction
The paper writes all FFN/MoE variants in one form, y(x)=A(Hact)∑igi(x)oi(x), where gi(x) specializes to an always-on block (dense FFN), all-experts-active normalized routing (Dense MoE), or top-a normalized routing (sparse MoE). Routing is restricted to token-choice (top-Ba/N0), which fixes the active set size deterministically per token; this makes the Ba/N1P update-size matching exact and avoids the train–inference mismatch of expert-choice routing.
Bridge I (dense FFN ↔ Dense MoE) applies active-width Ba/N2P at fixed backbone width Ba/N3: output multiplier Ba/N4, down-projection init std Ba/N5, and unchanged down-projection learning rate. Normalized routing would otherwise shrink the routed update by Ba/N6, so Complete-muE introduces a route scale Ba/N7. An optional finite-width correction Ba/N8 exists but is set to 1, justified by an expansion showing Ba/N9 stays order-one near uniform logits.
Bridge II (Dense MoE ↔ sparse MoE) handles changes in activated experts. After layer-level matching via B0, the remaining effect is stochastic: each expert's batch and duration both scale as B1. The first-order SDE LR/WD multiplier B2 therefore cancels exactly, leaving B3, B4. The paper is explicit that this is not a strict SDE invariance: the signal-to-noise parameter shifts as B5, which is not absorbed by the first-order correction, so mild bounded hyperparameter drift across B6 is expected theoretically and observed empirically. Larger B7 improves expert-side SNR and lowers attainable loss even at unchanged raw hyperparameters—an important consequence, since it means sparsity reduction yields quality gains without retuning.
All remaining cases are compositions rather than new primitives. Capacity scaling (B8 at fixed B9) decomposes into a Dense-MoE width step followed by reverse sparsification; the width factors cancel, recovering the same rule up to N0, again with mild drift from the non-strict Bridge II behavior. Granularity scaling at fixed density N1 reduces exactly to Dense-MoE width transfer because expert batch and duration (N2, N3) are invariant to the partition N4. Shared, group-balanced, and hybrid blocks are treated as one expanded FFN of total active width N5 with a single global route scale on routed groups only. Global AdamW factors compose separately: schedule changes (batch/duration) apply the standard N6 rule uniformly—exact when total tokens are fixed, approximate with a N7 shift when iterations are fixed—while MoE architectural changes require no additional global multiplier.
Empirical validation
The evaluation uses controlled proxy sweeps (LM: N8, 32 layers, 25k steps; diffusion transformer: flow matching, 100k steps) plus large-scale runs. Across activated-expert counts (64e2a–64e16a), capacity, granularity, shared experts, group-balanced routing, depth, width, and FFN expansion ratio, the low-loss regions for LR, weight decay, and init std remain broad and aligned, with representative stable ranges such as LR N9–Ha=ah0 and WD Ha=ah1–Ha=ah2 for LM. The most direct evidence for the recipe fixes AdamW hyperparameters at dense-tuned values and varies only the MoE architecture along four axes: loss decreases consistently for both LM and diffusion without per-setting retuning. Notably, increasing shared experts consistently raises loss, leading the authors to conclude that shared experts are unnecessary for quality unless needed for compute–communication overlap—a claim that runs against the prevalence of shared experts in recent production systems.
A single-H100 latency benchmark quantifies the scaling trade-off between capacity and granularity. Capacity scaling from 8 to 256 experts costs only Ha=ah3–Ha=ah4 over the dense baseline (87.8–97.0 ms vs. 81.0 ms/step), whereas granularity scaling raises latency from 84.1 to 135.0 ms/step, and dense width scaling reaches 667.6 ms/step before exhausting memory. Capacity is therefore the cheaper axis at large per-device token batches.
Large-scale runs apply one hyperparameter setting per modality family—four diffusion regimes (256P/512P images, 240P key frames, 240P 5s videos) share LR Ha=ah5, WD Ha=ah6, and the LM run uses LR Ha=ah7, WD Ha=ah8—with backbone width Ha=ah9 the proxy width (a0). At roughly 0.62B active / 6.29B total parameters and 100k steps, MoE achieves approximately a1 convergence speedup on 256P images, a2 on 240P 5s video, and a3–a4 on LLM training relative to dense during the stable-LR phase. Downstream benchmarks improve the 13-task average from 44.3 (dense) to 49.1 (128e8a4g1s) and 50.6 (128e8a1s); the non-grouped variant leads on average, consistent with smaller-scale findings that group-balanced routing slightly lags.
Limitations and open questions
The paper concedes several boundaries explicitly. First, Bridge II is not an exact invariance: the residual a5 shift produces bounded but nonzero hyperparameter drift across activated experts, capacity, and fixed-iteration batch changes, so the "near-optimal" transfer claim rests on empirical smallness of this drift rather than theory. Second, the cancellation argument assumes approximate load balancing; the imbalanced-routing extension shows the cancellation survives only at the expert-averaged level, with load fluctuations contributing to drift. Third, the route-scale rule discards the finite-width factor a6, which can deviate from one if selected router logits become highly separated during training—the analysis covers initialization but not fully trained routers. Fourth, the capacity–noise trade-off means increasing a7 need not monotonically improve loss; saturation or reversal is possible once the SNR penalty a8 dominates, and the paper does not characterize where this crossover occurs. Finally, the framework is restricted to token-choice routing and the AdamW gradient-magnitude-normalized regime; extension to expert-choice routing or other optimizers remains open.
Conclusion
Complete-muE provides a compositional AdamW transfer rule covering activated experts, total capacity, fixed-density granularity, shared/group-balanced hybrids, and standard width/depth/batch/duration changes, built from two bridges whose key insight is that expert-side batch and duration ratios cancel in the first-order SDE correction. Controlled sweeps and large-scale multimodal/LM runs support the central operational claim—that a single dense calibration suffices for near-optimal MoE training—with convergence speedups up to a9 over dense baselines at matched active parameter count. The residual theoretical gap between the demonstrated "relatively stable" transfer and strict invariance, and its dependence on load-balancing assumptions, remain the main open questions.