On-policy success conditioning and its sample complexity

Develop on-policy variants of success conditioning that estimate Q-functions from trajectory samples, and investigate the sample complexity of these variants.

Background

The success-conditioning update maintains positive probability on every action supported by the initial policy, which makes on-policy estimation of Q-functions conceptually possible through trajectory sampling. This property is presented as an advantage over deterministic value- and policy-iteration methods.

The paper identifies, in comparison with trust-region and proximal policy optimization, the design of trajectory-sample-based on-policy success-conditioning methods and the analysis of their sample complexity as concrete open directions.

References

For comparison to TRPO and PPO, interesting open directions include designing on-policy variants of SC that estimate the Q-functions from trajectory samples and investigating the sample complexity of these variants.

— On the Convergence of Success Conditioning for Policy Optimization  (2610.03642 - Brun et al., 2 Oct 2026) in Section 6, Discussion and Conclusion