Papers
Topics
Authors
Recent
Search
2000 character limit reached

ReaLB: Real-Time Load Balancing for Multimodal MoE Inference

Published 21 Apr 2026 in cs.DC | (2604.19503v1)

Abstract: Mixture-of-Experts (MoE) architectures are widely used in modern LLMs and multimodal models. However, inference efficiency is often limited by highly dynamic and skewed expert workloads across different modalities. During the prefill stage with large batch sizes, vision tokens frequently dominate the input sequences. Under expert parallelism (EP), this leads to severe load imbalance, where a subset of devices becomes overloaded, reducing overall system throughput. We propose ReaLB, a real-time load balancing method for multimodal MoE (MMoE) inference that introduces zero scheduling overhead. ReaLB dynamically adjusts the computation precision of MoE experts at runtime on a per-EP-rank basis. For ranks dominated by vision-heavy experts, ReaLB assigns lower-precision computation to improve execution efficiency by exploiting FP4 Tensor Cores. ReaLB does not require redundant experts or additional memory allocation. Instead, it performs layer-wise expert precision transformation on the fly and hides the associated overhead within the dispatch phase before MoE computation. Experiments on representative MMoE models show that ReaLB achieves 1.29x layer-level speedup while limiting accuracy loss to within 1.2%.

Summary

  • The paper introduces ReaLB, a real-time load-balancing system that detects hot, vision-heavy ranks and accelerates their experts with FP4 instead of migrating tokens or replicating weights.
  • ReaLB overlaps weight quantization with all-to-all dispatch and achieves up to 1.29× MoE-layer speedup and 1.06×–1.53× end-to-end throughput on multimodal MoE models.
  • The method avoids extra expert memory and keeps average accuracy within 1.2 points of BF16, but its benefits require large batches and remain less reliable for text-heavy workloads and larger server-scale systems.

ReaLB: Real-Time Load Balancing for Multimodal MoE Inference

Problem and motivation

Under expert parallelism (EP), Mixture-of-Experts (MoE) inference is globally synchronized across ranks, so any device that receives a disproportionate token load becomes a straggler and dictates layer latency. The paper argues that this imbalance is especially acute in multimodal MoE (MMoE) inference during large-batch prefill, where vision tokens dominate the workload and their routing distribution is both highly skewed and highly volatile across iterations.

The authors profile Kimi-VL under EP=8 on MMMU over 500 iterations and quantify three properties of this regime. First, device-level load skew is substantial: hot experts such as E4E_4 and E5E_5 on Rank0Rank_0 process 5–6× more tokens than the average expert, and the ratio of the most-loaded expert to the average fluctuates between 2× and 12× within short windows. Second, the imbalance is modality-driven: the vision-token fraction on the top-1 hot expert varies from 31% to 93% across ranks. Third, and critically for the design, inference routing lacks the temporal locality observed in MoE training—the identity of the hottest device and expert can change entirely between consecutive iteration windows. The paper shows concretely that EPLB, configured with a sliding window of 200 iterations and a 300-iteration rebalancing interval, mispredicts the hot spot (E33/Rank3E_{33}/Rank_3 versus E11/Rank1E_{11}/Rank_1), leaving residual stragglers.

This observation underpins a strong claim in the paper: prediction-based load balancing faces a fundamental difficulty in MMoE inference, because increasing prediction frequency does not improve accuracy under such dynamics while each invocation incurs non-trivial overhead. The paper also quantifies the costs of the replication-based alternative: migrating KK expert replicas costs K×SizeexpertK \times Size_{expert} in communication, and for DeepSeek-V3, one redundant expert per EP rank adds roughly 2.4 GB of memory, directly constraining achievable batch size.

Design

ReaLB rebalances load not by moving experts or tokens, but by changing the execution precision of experts on overloaded EP ranks at runtime. The mechanism exploits the fact that vision tokens exhibit higher redundancy along the forward pass than text tokens, so accelerating vision-heavy stragglers with FP4 Tensor Cores preserves accuracy where it matters most.

The system operates in four stages per MoE layer: (1) routing statistics are gathered after gating; (2) a scheduler classifies each rank as hot if Loadd/Ideald>CLoad_d / Ideal_d > C (capacity factor CC) and as vision-heavy if Rvd>MdR_{vd} > M_d (the vision-token ratio threshold); (3) a pipeline orchestrator performs online BF16-to-FP4 weight transformation overlapped with the all-to-all dispatch; (4) ranks execute with the assigned precision, with activation quantization fused into the expert GEMM kernels since it depends on received tokens.

Three design decisions are notable. Zero memory overhead: experts store only original high-precision weights plus precomputed scaling factors, avoiding both redundant replicas and multiple precision copies. Zero critical-path overhead: weight quantization runs concurrently with dispatch communication, and the paper claims this fully hides scheduling and transformation latency; a sequential variant (ReaLB-seq) is evaluated as an ablation. Compute-bound gating: ReaLB activates only above a global batch threshold (2048 tokens across 8 ranks), because below ~256 tokens per rank, MoE latency is dominated by non-GEMM operations and imbalance is irrelevant. The global threshold ensures all ranks trigger simultaneously, avoiding synchronization inconsistency.

For modality-isolated MMoEs such as ERNIE-VL, ReaLB is applied only to vision MoE layers, bypassing the modality threshold entirely.

Evaluation

ReaLB is implemented in vLLM v0.13.0 with NVFP4 GEMM kernels from FlashInfer, evaluated on Kimi-VL, Qwen3-VL-30B-A3B, and ERNIE-4.5-VL-27B-A3B on 8× RTX 5090. The headline results are:

Metric Result
MoE-layer speedup up to 1.29× (Qwen-VL)
End-to-end throughput 1.06×–1.53×
Average accuracy loss within 1.2 points of BF16 baseline

The rank-level analysis shows the mechanism works as intended: in Qwen-VL, FP4 acceleration of E5E_50 shifts the slowest rank to E5E_51, yielding a 1.30× layer speedup; in ERNIE-VL, four ranks are accelerated while others retain BF16, giving 1.26×. EPLB, by contrast, shows unstable speedups and even degrades performance in some cases (e.g., 0.81× on Qwen-VL), because history-based replacement can worsen actual imbalance. The end-to-end ablation is instructive: ReaLB-seq's much smaller gains confirm that pipeline overlap is essential, and Async_EPLB performs nearly identically to EPLB because expert weight transfer is difficult to hide.

The accuracy results support the central modality-aware claim. Uniform FP4 (FP4-All) is consistently harmful—dropping MMMU by 4.22 points on Kimi-VL and DynaMath by 8.19 points—whereas ReaLB with E5E_52 achieves zero loss on MMMU and substantially smaller losses on math benchmarks. A sensitivity study shows speedup is largely insensitive to E5E_53 (1.28×–1.30× across 0, 0.7, and 0.9) while accuracy improves monotonically with higher thresholds, though the trade-off is task-dependent (E5E_54 loses 0.66 on MMMU where E5E_55 loses nothing).

Limitations and open questions

The paper is candid about several constraints. The evaluation hardware is a consumer-grade cluster with limited interconnect bandwidth; end-to-end numbers substitute H20/NVLink communication latency measurements into the 5090 computation profile, so the throughput results are estimates rather than measurements on server-grade systems. The tested models are lightweight (A3B-class) relative to the stated target of models like Qwen3-VL-235B, and the claimed benefits at larger EP scale are argued rather than demonstrated. On RealWorldQA, ReaLB loses roughly 3 points—comparable to FP4-All—indicating that text tokens are not fully shielded from low-precision execution. The modality threshold E5E_56 is manually configured; adaptive tuning (e.g., AIMD-style) is proposed but not implemented. Latency spikes are observed for both ReaLB and EPLB, attributed respectively to modality-threshold filtering protecting text-heavy hot devices and to EPLB's coarse adjustment granularity. Finally, the design currently supports only FP4 as the lowest precision, leaving open how a multi-format precision ladder would be scheduled.

Conclusion

ReaLB reframes MoE load balancing as an execution-time problem solved by per-rank precision adaptation rather than workload prediction or expert migration. Its empirical contribution is demonstrating that modality-aware, on-the-fly FP4 switching—overlapped with dispatch communication—can deliver 1.29× layer-level and up to 1.53× end-to-end speedups on MMoE inference with bounded accuracy loss, without redundant experts or additional memory. The main open questions are whether the gains hold at server scale with high-bandwidth interconnects and larger models, and whether the accuracy protection can be extended to text-heavy workloads where the current mechanism offers no advantage.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.