Explore separate GRPO advantage groups for conditioned and unconditioned rollouts

Investigate whether separating insight-conditioned and unconditioned rollouts into distinct GRPO advantage groups produces benefits beyond the shared-group objective used in RLTL;DR.

Background

RLTL;DR collects sequential rollouts that may have different insight contexts but treats the attempts as a single GRPO group. An ablation found no performance difference when splitting the advantage groups, but the authors retained the simpler shared-group formulation.

The effect may therefore require further study under different training conditions, group constructions, or task distributions.

References

We leave further exploration of this to future work, since our primary goal is simplicity.

— RLTL;DR: Self-improvement by Internalizing Self-generated Feedback  (2609.37633 - Kirchhof et al., 29 Sep 2026) in Appendix, Section “Loss functions and GRPO advantage groups”