High-performance and cost-efficient datacenter network architecture for large-scale LLM training
Determine a datacenter network architecture that simultaneously delivers high performance and cost-efficiency for large-scale large language model training workloads.
References
In summary, how to design a high performance and cost-efficient datacenter network architecture for large-scale LLM training is still an open problem.
— UB-Mesh: a Hierarchically Localized nD-FullMesh Datacenter Network Architecture
(2503.20377 - Liao et al., 26 Mar 2025) in End of Subsection 2.3 "Datacenter Network Architectures" (Section 2)
Several aspects remain open. Our evaluation uses a limited set of models with fixed parameters, constraining analysis of compute-to-communication balance and achievable overlap. Our Alps/LUMI microbenchmarks suggest collective-library convergence behavior warrants further study to disentangle noise from genuine contention.
— Characterizing the Scalability and Performance of Large-Scale AI Training Under Multi-Tenancy
(2609.00817 - Raffi et al., 1 Sep 2026) in Section 6, “Conclusion and Future Work”
However, exploiting a THz overlay for distributed training remains an open problem.
— THz-SynC: Collective Synthesis with Contextual-Bandit-Assisted Coordination for Reconfigurable Hybrid Optical-THz AI Datacenters
(2609.04025 - Jiang et al., 3 Sep 2026) in Section I, Introduction