Scaling self-play pretraining to larger models

Determine whether the predictable transfer and scaling results obtained with Self-Play Pretraining with Zero Data persist when the self-play learner and generator are scaled beyond 25 million parameters while using contexts longer or comparable to the reported 4K-token context.

Background

The paper demonstrates zero-shot transfer from synthetic data generated by a self-play system in which a generator proposes programs for a universal Turing machine and a learner predicts the resulting byte sequences. The experiments are limited to models below 25 million parameters and a 4K-token context, and they report predictable improvements on held-out natural datasets across several modalities.

The authors identify persistence at larger model scales as the most immediate unresolved question. They note that addressing it may require a more expressive programming language with reusable abstractions that can co-evolve with the generator, as well as improvements to program-search efficiency and overall scalability. The problem is therefore to establish whether the observed transfer and scaling behavior survives beyond the current experimental regime.

References

The experiments reported here are confined to models below 25M parameters at a 4K context, and the most immediate question is whether these results persist for larger models.

— Self-Play Pretraining with Zero Data  (2609.30063 - Cowsik et al., 24 Sep 2026) in Section 6, paragraph “Future Directions”