Generalize SRT across scales, model families, and languages

Investigate whether Spaced Repetition Training maintains its observed behavior at larger model scales, across model families beyond the Llama family, and in multilingual settings.

Background

The language experiments evaluate only TinyLlama-1.1B-Chat and Llama-3.2-3B-Instruct, both from the Llama family, and use English data. The paper therefore explicitly identifies the behavior of Spaced Repetition Training outside these model-scale, architecture, and language settings as unresolved.

References

Our language experiments use TinyLlama-1.1B and Llama-3.2-3B-Instruct, both from the Llama family, and evaluation is restricted to English; behaviour at larger scales, across different model families, and in multilingual settings remains open, even though our cross-modality evidence demonstrates the effectiveness of SRT.

When to Review: Spaced Repetition for Continual Pre-Training of Language Models  (2608.17530 - Atreya et al., 18 Aug 2026) in Limitations section