Explore separate GRPO advantage groups for conditioned and unconditioned rollouts
Investigate whether separating insight-conditioned and unconditioned rollouts into distinct GRPO advantage groups produces benefits beyond the shared-group objective used in RLTL;DR.
References
We leave further exploration of this to future work, since our primary goal is simplicity.
— RLTL;DR: Self-improvement by Internalizing Self-generated Feedback
(2609.37633 - Kirchhof et al., 29 Sep 2026) in Appendix, Section “Loss functions and GRPO advantage groups”