Extension of Weave to Larger and Multi-Node Deployments

Extend the Weave runtime SM scheduling system for Mixture-of-Experts inference from a single 4×H100 NVLink node with expert parallelism degree 4 to larger GPU counts and multi-node configurations.

Background

Weave is evaluated only on a single node containing four NVIDIA H100 GPUs connected through NVLink, using expert parallelism with degree 4. Its runtime spatial and temporal schedulers adapt communication/computation SM allocation and execution order to per-layer routing results within this single-node setting.

The paper identifies deployment at larger GPU scales and across multiple nodes as unresolved future work. Such configurations would introduce broader communication topologies and potentially different communication bottlenecks, which may require extending Weave’s routing-aware scheduling and cost-model mechanisms beyond the validated four-GPU NVLink environment.

References

Due to hardware constraints, Weave is currently validated on a single 4$\times$H100 NVLink node with EP=4. Extending to larger GPU counts and multi-node configurations is left for future work.

— Weave: Fine-Grained Dynamic SM Scheduling in an MoE Megakernel for Compute-Communication Overlap  (2609.21483 - Huang et al., 18 Sep 2026) in Section 7, “Conclusion,” subsection “Limitations and future work”

Another unresolved question is how Flux can accommodate \gls{MoE} models.

— Flux: Optimal Scheduling of Optical Circuit Switches for LLM Training  (2609.25949 - Troch et al., 22 Sep 2026) in Section 4, subsection “Discussion”