Extension of Weave to Larger and Multi-Node Deployments
Extend the Weave runtime SM scheduling system for Mixture-of-Experts inference from a single 4×H100 NVLink node with expert parallelism degree 4 to larger GPU counts and multi-node configurations.
References
Due to hardware constraints, Weave is currently validated on a single 4$\times$H100 NVLink node with EP=4. Extending to larger GPU counts and multi-node configurations is left for future work.
— Weave: Fine-Grained Dynamic SM Scheduling in an MoE Megakernel for Compute-Communication Overlap
(2609.21483 - Huang et al., 18 Sep 2026) in Section 7, “Conclusion,” subsection “Limitations and future work”
Another unresolved question is how Flux can accommodate \gls{MoE} models.
— Flux: Optimal Scheduling of Optical Circuit Switches for LLM Training
(2609.25949 - Troch et al., 22 Sep 2026) in Section 4, subsection “Discussion”