Papers
Topics
Authors
Recent
Search
2000 character limit reached

Fine-Tuning on Self-Generated and Reward-Weighted Data: Learning Dynamics, Convergence Rates, and Benefits of Off-Policyness

Published 29 Sep 2026 in cs.LG and math.OC | (2609.36945v1)

Abstract: We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version of REINFORCE -- referred to as RE(S) -- that updates the rollout distribution once every S≥1S \ge 1 gradient steps. Prior work in bandits and reinforcement learning has developed rich theory for policy gradient methods, and on-policy sampling (i.e., a small SS, ideally $1$) is often viewed as crucial to their success; yet in prominent application like post-training LLMs, reward-guided self-training has proved to be effective even when the rollout distribution is updated infrequently, but theoretical understanding remains limited for the convergence properties of these off-policy methods. To bridge these gaps, we develop a unified theory for RE(S) that covers the full spectrum of S≥1S \ge 1: it can be interpreted as a stage-wise optimization process, where each stage takes SS gradient steps for minimizing the Kullback-Leibler distance to a fixed reward-weighted rollout distribution. For multi-arm bandits with softmax policies, our in-depth analysis and numerical experiments reveal three key findings: (1) for any fixed SS, RE(S) enjoys global convergence to the optimal policy as the number of rollout distribution updates B=⌊T/S⌋→∞B = \lfloor T / S \rfloor \rightarrow \infty, where TT denotes the number of gradient steps; (2) we prove tight two-sided bounds showing that the suboptimality gap of RE(S) achieves an asymptotic Θ(1/T)Θ(1 / T) convergence rate, while SS only affects the length of a burn-in phase; (3) when initialized at a weak policy with a small optimal-action probability, RE(1) gets trapped around suboptimal policies for a long period, whereas RE(S) with a suitable SS avoids the detour and achieves significantly faster convergence to the global optimum, highlighting the benefits of off-policyness in this case.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.