Does extended RL fine-tuning (PPO/GRPO) cause model collapse?

Determine whether longer training periods using reinforcement learning techniques such as Proximal Policy Optimization (PPO) or Group Relative Optimization (GRPO) lead to model collapse in large language models.

Background

The authors relate iterative deployment to reinforcement learning and discuss safety implications of implicit reward signals arising from user curation. In this context, they raise uncertainty about the stability of models under extended RL fine-tuning regimes commonly used to improve reasoning skills.

Specifically, they note that it is unclear whether prolonged RL training—using methods like PPO or GRPO—induces model collapse, which would manifest as degraded capabilities due to distributional shrinkage.

References

Note, however, that it is also unclear whether longer training periods using RL techniques, such as PPO or GRPO, lead to model collapse or not.

Iterative Deployment Improves Planning Skills in LLMs  (2512.24940 - Corrêa et al., 31 Dec 2025) in Subsection: Implications to AI Safety

The original collapse endpoint is unresolved. Extending four arms to $300$ steps---ThinkPrior and online-NP, each with and without a reference KL at $\beta{=}0.04$, $20$ runs---the original absolute endpoint gives online-NP $2/6$ against ThinkPrior $0/6$ in the KL-free pair, Fisher two-sided $p{=}0.455$. After inspecting trajectories, a relative-drop threshold changes that count to $4/6$ against $0/6$, $p{=}0.061$; this revised endpoint is exploratory, post hoc and still not significant. It is an observation to replicate, not a stability result or a causal attribution to the selection rule (Appendix~D).

ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR  (2609.09075 - Sha et al., 8 Sep 2026) in Section Analysis, paragraph “The original collapse endpoint is unresolved”; Appendix D, “Original endpoint” and “What we do and do not claim”

We do not claim that a majority-vote reward does this in general: seed~$2$ took two intact updates at the identical learning rate, so the rate is not uniformly fatal, and the band arm at a shorter cap on an easier stream never collapses at all. Separating the reward from the step size needs a learning-rate sweep we did not run.

Phantom Gains: Auditing Self-Improvement Against a Measured Null  (2608.20290 - Xu et al., 20 Aug 2026) in Appendix, Section “Policy gradient on a majority-vote reward,” paragraph “The scope of the claim, and the exclusion criterion”