Papers
Topics
Authors
Recent
Search
2000 character limit reached

On the Convergence of Success Conditioning for Policy Optimization

Published 2 Oct 2026 in cs.LG and math.OC | (2610.03642v1)

Abstract: Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common to many reinforcement learning applications, yet its limiting behavior and convergence rates are not well understood. In this work, we demonstrate that success conditioning converges to an optimal policy on a broad class of Markov decision processes (MDPs). We also derive convergence rates in some common settings. For discounted MDPs, we prove convergence within O(1/ε<sup>p)\mathcal{O}(1/\varepsilon<sup>p) iterations to an ε\varepsilon-optimal policy, where the exponent pp depends on problem data. For single-period MDPs, such a policy is obtained within O(log⁡(1/ε))\mathcal{O}(\log(1/\varepsilon)) iterations.

Authors (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.