- The paper introduces ReaLB, a real-time load-balancing system that detects hot, vision-heavy ranks and accelerates their experts with FP4 instead of migrating tokens or replicating weights.
- ReaLB overlaps weight quantization with all-to-all dispatch and achieves up to 1.29× MoE-layer speedup and 1.06×–1.53× end-to-end throughput on multimodal MoE models.
- The method avoids extra expert memory and keeps average accuracy within 1.2 points of BF16, but its benefits require large batches and remain less reliable for text-heavy workloads and larger server-scale systems.
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
Problem and motivation
Under expert parallelism (EP), Mixture-of-Experts (MoE) inference is globally synchronized across ranks, so any device that receives a disproportionate token load becomes a straggler and dictates layer latency. The paper argues that this imbalance is especially acute in multimodal MoE (MMoE) inference during large-batch prefill, where vision tokens dominate the workload and their routing distribution is both highly skewed and highly volatile across iterations.
The authors profile Kimi-VL under EP=8 on MMMU over 500 iterations and quantify three properties of this regime. First, device-level load skew is substantial: hot experts such as E4 and E5 on Rank0 process 5–6× more tokens than the average expert, and the ratio of the most-loaded expert to the average fluctuates between 2× and 12× within short windows. Second, the imbalance is modality-driven: the vision-token fraction on the top-1 hot expert varies from 31% to 93% across ranks. Third, and critically for the design, inference routing lacks the temporal locality observed in MoE training—the identity of the hottest device and expert can change entirely between consecutive iteration windows. The paper shows concretely that EPLB, configured with a sliding window of 200 iterations and a 300-iteration rebalancing interval, mispredicts the hot spot (E33/Rank3 versus E11/Rank1), leaving residual stragglers.
This observation underpins a strong claim in the paper: prediction-based load balancing faces a fundamental difficulty in MMoE inference, because increasing prediction frequency does not improve accuracy under such dynamics while each invocation incurs non-trivial overhead. The paper also quantifies the costs of the replication-based alternative: migrating K expert replicas costs K×Sizeexpert in communication, and for DeepSeek-V3, one redundant expert per EP rank adds roughly 2.4 GB of memory, directly constraining achievable batch size.
Design
ReaLB rebalances load not by moving experts or tokens, but by changing the execution precision of experts on overloaded EP ranks at runtime. The mechanism exploits the fact that vision tokens exhibit higher redundancy along the forward pass than text tokens, so accelerating vision-heavy stragglers with FP4 Tensor Cores preserves accuracy where it matters most.
The system operates in four stages per MoE layer: (1) routing statistics are gathered after gating; (2) a scheduler classifies each rank as hot if Loadd/Ideald>C (capacity factor C) and as vision-heavy if Rvd>Md (the vision-token ratio threshold); (3) a pipeline orchestrator performs online BF16-to-FP4 weight transformation overlapped with the all-to-all dispatch; (4) ranks execute with the assigned precision, with activation quantization fused into the expert GEMM kernels since it depends on received tokens.
Three design decisions are notable. Zero memory overhead: experts store only original high-precision weights plus precomputed scaling factors, avoiding both redundant replicas and multiple precision copies. Zero critical-path overhead: weight quantization runs concurrently with dispatch communication, and the paper claims this fully hides scheduling and transformation latency; a sequential variant (ReaLB-seq) is evaluated as an ablation. Compute-bound gating: ReaLB activates only above a global batch threshold (2048 tokens across 8 ranks), because below ~256 tokens per rank, MoE latency is dominated by non-GEMM operations and imbalance is irrelevant. The global threshold ensures all ranks trigger simultaneously, avoiding synchronization inconsistency.
For modality-isolated MMoEs such as ERNIE-VL, ReaLB is applied only to vision MoE layers, bypassing the modality threshold entirely.
Evaluation
ReaLB is implemented in vLLM v0.13.0 with NVFP4 GEMM kernels from FlashInfer, evaluated on Kimi-VL, Qwen3-VL-30B-A3B, and ERNIE-4.5-VL-27B-A3B on 8× RTX 5090. The headline results are:
| Metric |
Result |
| MoE-layer speedup |
up to 1.29× (Qwen-VL) |
| End-to-end throughput |
1.06×–1.53× |
| Average accuracy loss |
within 1.2 points of BF16 baseline |
The rank-level analysis shows the mechanism works as intended: in Qwen-VL, FP4 acceleration of E50 shifts the slowest rank to E51, yielding a 1.30× layer speedup; in ERNIE-VL, four ranks are accelerated while others retain BF16, giving 1.26×. EPLB, by contrast, shows unstable speedups and even degrades performance in some cases (e.g., 0.81× on Qwen-VL), because history-based replacement can worsen actual imbalance. The end-to-end ablation is instructive: ReaLB-seq's much smaller gains confirm that pipeline overlap is essential, and Async_EPLB performs nearly identically to EPLB because expert weight transfer is difficult to hide.
The accuracy results support the central modality-aware claim. Uniform FP4 (FP4-All) is consistently harmful—dropping MMMU by 4.22 points on Kimi-VL and DynaMath by 8.19 points—whereas ReaLB with E52 achieves zero loss on MMMU and substantially smaller losses on math benchmarks. A sensitivity study shows speedup is largely insensitive to E53 (1.28×–1.30× across 0, 0.7, and 0.9) while accuracy improves monotonically with higher thresholds, though the trade-off is task-dependent (E54 loses 0.66 on MMMU where E55 loses nothing).
Limitations and open questions
The paper is candid about several constraints. The evaluation hardware is a consumer-grade cluster with limited interconnect bandwidth; end-to-end numbers substitute H20/NVLink communication latency measurements into the 5090 computation profile, so the throughput results are estimates rather than measurements on server-grade systems. The tested models are lightweight (A3B-class) relative to the stated target of models like Qwen3-VL-235B, and the claimed benefits at larger EP scale are argued rather than demonstrated. On RealWorldQA, ReaLB loses roughly 3 points—comparable to FP4-All—indicating that text tokens are not fully shielded from low-precision execution. The modality threshold E56 is manually configured; adaptive tuning (e.g., AIMD-style) is proposed but not implemented. Latency spikes are observed for both ReaLB and EPLB, attributed respectively to modality-threshold filtering protecting text-heavy hot devices and to EPLB's coarse adjustment granularity. Finally, the design currently supports only FP4 as the lowest precision, leaving open how a multi-format precision ladder would be scheduled.
Conclusion
ReaLB reframes MoE load balancing as an execution-time problem solved by per-rank precision adaptation rather than workload prediction or expert migration. Its empirical contribution is demonstrating that modality-aware, on-the-fly FP4 switching—overlapped with dispatch communication—can deliver 1.29× layer-level and up to 1.53× end-to-end speedups on MMoE inference with bounded accuracy loss, without redundant experts or additional memory. The main open questions are whether the gains hold at server scale with high-bandwidth interconnects and larger models, and whether the accuracy protection can be extended to text-heavy workloads where the current mechanism offers no advantage.