原因 of degraded reuse performance in Hopper-v5 at large windows

Investigate whether the degraded training of uniform and balance-heuristic-corrected sample-reuse variants of Proximal Policy Optimization in Hopper-v5 at reuse-window size \(\omega=8\) is caused by the combined effects of stale advantages accumulated over many additional updates, the small fraction of fresh samples in each mini-batch, and increased drift among consecutive policies.

Background

The paper reports that, under the fixed mini-batch-size setting, the ω=4\omega=4 sample-reuse variants generally outperform vanilla Proximal Policy Optimization, whereas Hopper-v5 degrades when the reuse window is increased to ω=8\omega=8. The authors identify several possible interacting causes: substantially more actor updates, dilution of fresh on-policy data, and greater policy divergence that can push reused samples outside the clipping range.

Because the passage presents these causes as a conjecture rather than establishing which mechanism is responsible, determining the source of the Hopper-v5 failure remains an explicit unresolved problem.

References

We conjecture this results from several compounding factors, amplified by this environment's configuration ($H = 2048$, $B = 64$, $\widetilde{K} = 10$). First, at $\omega = 8$ the reuse variants perform $2560$ updates per iteration against the $320$ of \ppo, potentially amplifying the bias introduced by stale advantages across many gradient steps. Second, with $B/\omega = 8$ fresh samples per mini-batch, the fraction of on-policy signal per update is particularly small. Third, since all hyperparameters are inherited from \ppo, the larger number of updates may increase the drift between consecutive policies and hence the diversity of the policies in the window, possibly making past samples falling outside the clipping range more frequently.

— Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?  (2610.01399 - Montenegro et al., 1 Oct 2026) in Section 4, subsection “Fixed $B$”

We conjecture that the primary cause is the severely reduced amount of fresh data available to the critic. Since the critic is updated exclusively on the most recent dataset $\cD_k$, reducing new interactions to $H/\omega$ may directly hinder it, possibly leading to degraded advantage estimates that in turn corrupt actor updates. We anticipate that this is not mitigated by fitting the critic on the whole $\cD(\omega_{k})$ (see Appendix~\ref{apx:exp-on-off}).

— Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?  (2610.01399 - Montenegro et al., 1 Oct 2026) in Appendix, Section “Additional Experiments,” subsection “Reduced Data Collection Setting”