Scaling L0-MoE to larger language models
Determine whether scaling the L0-MoE architecture to larger language models, including models with approximately 70 billion parameters, can achieve greater inference speedups while maintaining performance comparable to the corresponding dense models.
References
However, due to computational resource constraints, we have not yet experimented with larger models (e.g., 70B parameters). We hypothesize that larger LLMs could potentially achieve even greater speedups. We leave the verification of this hypothesis for larger-scale models as future work.
— Accelerating Dense LLMs via L0-regularized Mixture-of-Experts
(2609.21672 - Zhang et al., 18 Sep 2026) in Section 4, Discussion; Section 5, Conclusion and Future Work