Papers
Topics
Authors
Recent
Search
2000 character limit reached

UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing

Published 2 Jun 2026 in cs.DC and cs.LG | (2606.04101v1)

Abstract: Large-scale expert parallelism (EP) is becoming pivotal for training and serving frontier MoE models, but it also amplifies device-level expert load imbalance into compute stragglers, token all-to-all bottlenecks, and activation-memory spikes. Existing balancers redistribute experts periodically based on historical load, which becomes unreliable for production deployments with non-stationary load patterns. We present UltraEP, the first exact-load, real-time balancer for large-EP MoE training and serving prefill on rack-scale nodes (RSNs). Built upon the extended scale-up connectivity of RSNs, UltraEP rebalances every microbatch and layer on critical paths, which requires nontrivial co-design of plan solving and expert replication communication to minimize exposed overhead. To this end, UltraEP eagerly reacts to post-gating load with efficient quota-driven planning, and executes the resulting irregular expert-state transfers with RSN-native persistent tile streaming and relay-based fan-out mitigation. Averaged across MoE models from 106B to 671B parameters in training and prefill, UltraEP achieves 94.3% of the force-balanced ideal throughput, delivering 1.49×\times improvement over non-balancing, while reducing the final inter-rank imbalance from 1.30−-4.01 to 1.01−-1.04. Additionally, we validate UltraEP's scalability and robustness in production MoE training with 2560 GPUs.

Summary

  • The paper introduces UltraEP, a GPU-based quota solver and RSN-native communication system that jointly replicates and reroutes experts after every microbatch and layer.
  • UltraEP reaches 94.3% of force-balanced ideal throughput, improves throughput by 1.49× over no balancing, and reduces inter-rank imbalance to 1.01–1.04 across 106B–671B models.
  • The system reduces activation-memory spikes and scales to 2,560 production GPUs, but depends on rack-scale bandwidth and does not address decode-phase or token-level routing imbalance.

Large-scale expert parallelism (EP) is the dominant scaling axis for frontier Mixture-of-Experts (MoE) models, but it amplifies device-level expert load imbalance into compute stragglers, token all-to-all bottlenecks, and activation-memory spikes on overloaded ranks. Existing balancers such as EPLB redistribute experts periodically using historical load statistics, an approach whose effectiveness depends on load stationarity. UltraEP departs from this paradigm: it is an exact-load, real-time balancer for large-EP MoE training and serving prefill on rack-scale nodes (RSNs), rebalancing every microbatch and every MoE layer on the critical path. Averaged across models from 106B to 671B parameters, it achieves 94.3% of the force-balanced ideal throughput, a 1.49× improvement over no balancing, while reducing final inter-rank imbalance from 1.30–4.01 down to 1.01–1.04, and it is validated in production training on 2560 GPUs.

Motivation: non-stationary expert load

The paper's empirical characterization motivates the shift from prediction-based to exact-load balancing. In serving prefill, expert popularity shifts sharply across semantic transitions (coding, science, mixed-domain traffic), across batches, and across layers; mixed-domain inputs superimpose multiple routing patterns and make imbalance less predictable. In training, early-phase router specialization is unstable, and even late-stage training retains strong dynamics: auxiliary-loss negative feedback continually re-adjusts expert utilization, and DeepSeek-style router compensation, while reducing average skew, actually enlarges short-term load swings. Inter-microbatch sampling jitter compounds this at fine granularity.

Under large-EP (e.g., 64-way), only two to four experts reside per rank, so routing dynamics translate directly into inter-rank skew. The paper shows that EPLB, rebalancing every 50 prefill steps or 3 global batches, cannot track these shifts; when realized load deviates from the statistics underlying the layout, EPLB can worsen imbalance and create new stragglers. This is a strong claim against the dominant deployed approach, and it motivates reacting to post-gating exact load rather than extrapolating stale measurements.

The design space is enabled by RSNs, which extend the scale-up domain from a single 4/8-GPU server to 64+ GPUs per rack with hundreds of GB/s to ~TB/s per-GPU bandwidth and load/store memory semantics. An entire EP group fits on one rack, making hot-path expert-state transfer physically viable — but RSNs are described as necessary, not sufficient. Two challenges remain: the control plane must produce a high-quality plan within the short gating-to-dispatch window, and the data plane must execute irregular, volatile transfers that static collective-oriented stacks handle poorly.

System design

UltraEP uses a fixed expert layout: each rank reserves main slots (hosting original expert instances, with full weight/gradient/optimizer buffers) and redundant slots (hosting replicas, with no optimizer state). It adopts replication-only balancing — main experts are never reordered — arguing that at large-EP the local main-expert set is too small for reordering to pay off. A key memory optimization is cross-layer buffer reuse for redundant slots: weights and gradients are shared across layers, shrinking a single redundant slot in Qwen3-235B from 3.3 GB weights / 6.6 GB gradients to 36 MB / 72 MB per rank, at the cost of a tight per-layer weight-materialization deadline on the forward critical path.

The forward pass reuses the existing notify-dispatch all-to-all to gather exact global load; every rank then deterministically computes an identical replication-and-reroute plan with no extra synchronization, distributes main-expert weights to replicas, reroutes tokens to physical instances, and dispatches. Planning and weight replication are on the critical path. The backward pass re-materializes replica weights (overlappable with Wgrad until Dgrad), then reduces replica gradients into main-expert buffers before the next layer, preserving exact training equivalence with the no-replica formulation. Backward reuses cached forward metadata, so only token all-to-all and MoE compute remain exposed.

Quota-driven planning

The core control-plane contribution is a joint replication–reroute solver driven by a per-instance load quota UU, where each quota specifies the final token load assigned to an expert instance. Rather than placing replicas first and patching with a separate reroute heuristic (as EPLB plus round-robin reroute does), UltraEP directly optimizes the post-reroute load. The solver binary-searches the smallest load threshold τ\tau such that every rank can be brought below it via replication alone, with a greedy feasibility oracle per probe that transfers load from the hottest experts of overloaded ranks to the admissible rank with the largest slack, subject to a quota floor umin⁡u_{\min}. Because a replica is materialized only when it carries useful load, the placement already encodes effective reroute capacity — avoiding ineffective replicas that satisfy placement heuristics but receive little traffic. Reroute then decomposes quotas source-wise with locality: local tokens consume the host rank's quota first, reducing cross-rank traffic without altering the solved threshold, and per-token assignment reduces to a prefix-scan lookup.

The solver runs fully on GPU — one cooperative thread block on a single SM, staging the load matrix and placement state in shared memory with warp-level parallelism — avoiding host synchronization on the hot path. Measured solving latency is 0.111 ms on average, 27.4% faster than EPLB+.

RSN-native communication

The data plane addresses two bottlenecks. First, persistent tile streaming: expert weights and gradients are divided into fixed-size tiles compiled into device-resident transfer tasks; a persistent kernel with double-buffered shared-memory tiles hides task lookup, address translation, and synchronization overhead. The thread-block footprint is an explicit resource knob — maximized on the forward critical path, bounded during overlapped backward to preserve SM headroom. Second, chunk-streaming relay for hot-expert fan-out: for experts whose replica count exceeds a threshold, the source seeds a relay set sized near ∣H(e)∣−1\sqrt{|\mathcal{H}(e)|-1}, which approximately balances the two forwarding stages; relays forward chunk-by-chunk as tiles arrive, pipelining the stages without global barriers. Relay topology is chosen load-aware, assigning relays and leaves to ranks with the smallest projected sending volume. Under high fan-out, relay provides an additional 1.3×–1.8× gain, sustaining near-constant ~0.28 ms latency while the no-relay variant grows linearly. Overall, UltraEP accelerates expert replication 3.1×–5.5× over torch.distributed and a DeepEP adaptation under identical plans.

Evaluation

Training experiments resume late-stage checkpoints of GLM4.5-106B, Qwen3-235B, and DeepSeek-V3 and run 20 global batches. Averaged across models, EPLB, LPLB, EPLB+, and UltraEP improve throughput over Megatron-LM by 20%, 12%, 29%, and 42% respectively. Notably, on DeepSeek-V3 — where router compensation enlarges short-term swings — EPLB and LPLB perform at or worse than no balancing, while UltraEP stays above 96% of ideal. In serving prefill, UltraEP reaches 90%–97% of ideal, delivering 1.56× and 1.29× the throughput of SGLang and EPLB respectively, and 5%–24% over EPLB+ (exact load with round-robin reroute, same communication), isolating the benefit of quota-driven planning.

The latency breakdown for Qwen3-235B shows hot-path overhead of only 0.33 ms in forward (1.8% of total) versus ideal; the residual gap is a 33% (forward) and 10% (backward) increase in token all-to-all latency, attributed to uneven token-level routing that the synthetic uniform-dispatch ideal does not exhibit — a candid attribution of the remaining gap to realistic routing irregularity rather than system overhead. Balancing also flattens receive-side activation hot spots: without balancing, MoE activation peak memory is 2× (training) and 11× (serving) the ideal, and UltraEP restores it to near-ideal, reducing OOM risk and potentially avoiding activation-checkpointing overhead.

In ablations, UltraEP achieves 1.03 average result imbalance versus 1.19 for EPLB+, using 57.9% fewer redundant slots and 3.9% less token traffic via locality. In production, 2560-GPU training of RefMoE-288B on 15T tokens sustains over 92% of force-balanced throughput with a 9.6% average gain over no balancing, and the loss curve follows the expected pretraining trajectory, confirming semantic preservation at scale (1.5M GPU hours).

Limitations and open questions

The paper is candid about several boundaries. The residual throughput gap to the force-balanced ideal stems from uneven token routing rather than rank-level imbalance, and the token all-to-all penalty (up to 33% in forward) is inherent to realistic routing — UltraEP does not address token-level irregularity within DeepEP's dispatch pipelines. The replication-only design forgoes main-expert reordering, justified empirically at large-EP but untested at smaller EP degrees where reordering may matter more. Cross-layer buffer reuse imposes a strict per-layer materialization deadline, making the approach sensitive to RSN scale-up bandwidth characteristics; the evaluation uses a single RSN hardware profile (64 GPUs, 900 GB/s intra-rack), and portability to RSNs with different memory-semantics or multicast capabilities (e.g., switch-offloaded multicast) is not examined. Decode-phase balancing is explicitly out of scope, on the argument that compute imbalance is diluted by memory-boundness under TPOT SLOs — an assumption that may weaken as batch sizes grow. Finally, the relay threshold and umin⁡u_{\min} are tuned hyperparameters whose sensitivity is not systematically characterized. The paper leaves open whether exact-load balancing extends to RL pipelines alternating training and inference phases, and how the quota abstraction composes with FP8/low-precision expert-state transfers.

Conclusion

UltraEP demonstrates that exact-load, real-time expert balancing on the hot path is feasible when an EP group is contained within an RSN scale-up domain and planning and communication are co-designed. Its quota-driven joint replication–reroute solver and RSN-native tile-streaming and relay communication yield 94.3% of ideal throughput, near-optimal 1.01–1.04 inter-rank imbalance, and validated scalability to 2560-GPU production training. The result establishes exact-load balancing as a practical replacement for history-based periodic rebalancing in large-EP deployments, contingent on rack-scale connectivity.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.