Source of diminishing returns from larger reuse windows

Determine whether the diminishing benefits of increasing the sample-reuse window in Proximal Policy Optimization arise from the increasing proportion of stale advantages and samples collected under increasingly distant behavioral policies within each mini-batch.

Background

In the fixed-mini-batch-count setting, the paper finds that sample reuse improves Proximal Policy Optimization across the evaluated environments, with most benefits appearing at ω=4\omega=4. Increasing the window to ω=8\omega=8, however, produces diminishing returns and only marginal improvements in several cases.

The authors conjecture that the larger share of stale advantages and data generated by more distant policies offsets the benefit of adding historical samples. The claim is explicitly presented as a conjectural explanation rather than a demonstrated causal result.

References

We conjecture this is due to the increasing proportion of stale advantages and of samples collected under increasingly distant policies within each mini-batch, which gradually offsets the benefit of additional data as the window grows.

— Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?  (2610.01399 - Montenegro et al., 1 Oct 2026) in Section 4, subsection “Fixed $N_{\cB}$”

We conjecture this is due to the larger policy changes occurring during early training: consecutive policies differ the most in this phase, so that, with a large reuse window, older samples carry the most stale advantages and come from the most distant behavioral policies.

— Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?  (2610.01399 - Montenegro et al., 1 Oct 2026) in Appendix, Section “Additional Experiments,” subsection “Fixed Mini-Batch Count Setting”