Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sample complexity of variance-reduced policy gradient: weaker assumptions and lower bounds

Published 2 Oct 2026 in cs.LG | (2610.03165v1)

Abstract: Several variance-reduced versions of REINFORCE based on importance sampling achieve an improved O(ε<sup>−3)O(ε<sup>{-3}) sample complexity to find an εε-stationary point, under an unrealistic assumption on the variance of the importance weights. In this paper, we propose the \algo (Defensive Policy Gradient) algorithm, based on defensive importance sampling, which achieves the same rate without any assumption on the variance of ordinary importance weights. We also establish lower bounds in a generalized black-box policy-optimization model that hides states and actions and permits parameter-dependent rewards. In this model, the optimal rates are Θ(ε<sup>−4)Θ(ε<sup>{-4}) with bounded-variance one-policy feedback and Θ(ε<sup>−3)Θ(ε<sup>{-3}) with mean-square-smooth coupled two-policy feedback. Under standard policy-regularity conditions, REINFORCE and \algo realize the corresponding oracle conditions and attain the O(ε<sup>−4)O(ε<sup>{-4}) and O(ε<sup>−3)O(ε<sup>{-3}) upper bounds, respectively. Although the lower bounds do not apply directly to the classical MDP interaction model in which these algorithms operate, this correspondence provides oracle-level evidence that the faster rate of \algo is optimal and genuinely separated from that of vanilla policy gradient.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.