Assess the stability of uniform-replay reasoning degradation

Determine whether the degradation of BBH and GSM8K reasoning performance observed under uniform replay remains stable across multiple independent training runs.

Background

At the Llama-3.2-3B-Instruct scale, uniform replay substantially degrades reasoning benchmarks, including BBH and GSM8K. Although the pattern is consistent across two benchmarks, each condition was trained only once, so the paper leaves unresolved whether the effect persists across independent training runs.

References

While the reasoning-benchmark degradation under uniform replay is consistent across two independent benchmarks (BBH and GSM8K) and is not exhibited by SRT under identical conditions, which makes a single-run artifact less likely, confirming its stability would require multiple independent training runs. We leave this to future work.

When to Review: Spaced Repetition for Continual Pre-Training of Language Models  (2608.17530 - Atreya et al., 18 Aug 2026) in Limitations section