On-policy success conditioning and its sample complexity
Develop on-policy variants of success conditioning that estimate Q-functions from trajectory samples, and investigate the sample complexity of these variants.
References
For comparison to TRPO and PPO, interesting open directions include designing on-policy variants of SC that estimate the Q-functions from trajectory samples and investigating the sample complexity of these variants.
— On the Convergence of Success Conditioning for Policy Optimization
(2610.03642 - Brun et al., 2 Oct 2026) in Section 6, Discussion and Conclusion