Does extended RL fine-tuning (PPO/GRPO) cause model collapse?
Determine whether longer training periods using reinforcement learning techniques such as Proximal Policy Optimization (PPO) or Group Relative Optimization (GRPO) lead to model collapse in large language models.
References
Note, however, that it is also unclear whether longer training periods using RL techniques, such as PPO or GRPO, lead to model collapse or not.
The original collapse endpoint is unresolved. Extending four arms to $300$ steps---ThinkPrior and online-NP, each with and without a reference KL at $\beta{=}0.04$, $20$ runs---the original absolute endpoint gives online-NP $2/6$ against ThinkPrior $0/6$ in the KL-free pair, Fisher two-sided $p{=}0.455$. After inspecting trajectories, a relative-drop threshold changes that count to $4/6$ against $0/6$, $p{=}0.061$; this revised endpoint is exploratory, post hoc and still not significant. It is an observation to replicate, not a stability result or a causal attribution to the selection rule (Appendix~D).
We do not claim that a majority-vote reward does this in general: seed~$2$ took two intact updates at the identical learning rate, so the rate is not uniformly fatal, and the band arm at a shorter cap on an easier stream never collapses at all. Separating the reward from the step size needs a learning-rate sweep we did not run.