Determine whether data repetition increases MoE load-balancing loss

Determine whether data repetition itself increases Mixture-of-Experts router load-balancing loss, independently of the optimization effects caused by extremely low training loss at high repetition rates.

Background

The paper analyzes router load-balancing loss across dense and Mixture-of-Experts LLMs trained with different data-repetition rates. At smaller model scales, the authors observe no consistent pattern, whereas at the 1B-active-parameter scale some intermediate and very high repetition rates produce unusual or periodic behavior. The authors therefore leave unresolved whether repetition directly increases the auxiliary load-balancing loss or whether the observed behavior is instead mediated by the very low training loss reached at high repetition rates, which may strengthen the optimization signal from auxiliary losses.

References

It is possible that repetition itself increases load balancing loss, but that the extremely low training loss at high $R$ results in a relatively strong optimization signal from auxiliary losses, eventually driving load balancing loss to fall.

Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data  (2609.11917 - Jha et al., 10 Sep 2026) in Appendix, Section “Additional Results,” subsection “Routing Load Balance and Stability,” caption of the routing load-balancing-loss figure