Scaling self-play pretraining to larger models
Determine whether the predictable transfer and scaling results obtained with Self-Play Pretraining with Zero Data persist when the self-play learner and generator are scaled beyond 25 million parameters while using contexts longer or comparable to the reported 4K-token context.
References
The experiments reported here are confined to models below 25M parameters at a 4K context, and the most immediate question is whether these results persist for larger models.
— Self-Play Pretraining with Zero Data
(2609.30063 - Cowsik et al., 24 Sep 2026) in Section 6, paragraph “Future Directions”