Assess the stability of uniform-replay reasoning degradation
Determine whether the degradation of BBH and GSM8K reasoning performance observed under uniform replay remains stable across multiple independent training runs.
References
While the reasoning-benchmark degradation under uniform replay is consistent across two independent benchmarks (BBH and GSM8K) and is not exhibited by SRT under identical conditions, which makes a single-run artifact less likely, confirming its stability would require multiple independent training runs. We leave this to future work.
— When to Review: Spaced Repetition for Continual Pre-Training of Language Models
(2608.17530 - Atreya et al., 18 Aug 2026) in Limitations section