原因 of degraded reuse performance in Hopper-v5 at large windows
Investigate whether the degraded training of uniform and balance-heuristic-corrected sample-reuse variants of Proximal Policy Optimization in Hopper-v5 at reuse-window size \(\omega=8\) is caused by the combined effects of stale advantages accumulated over many additional updates, the small fraction of fresh samples in each mini-batch, and increased drift among consecutive policies.
References
We conjecture this results from several compounding factors, amplified by this environment's configuration ($H = 2048$, $B = 64$, $\widetilde{K} = 10$). First, at $\omega = 8$ the reuse variants perform $2560$ updates per iteration against the $320$ of \ppo, potentially amplifying the bias introduced by stale advantages across many gradient steps. Second, with $B/\omega = 8$ fresh samples per mini-batch, the fraction of on-policy signal per update is particularly small. Third, since all hyperparameters are inherited from \ppo, the larger number of updates may increase the drift between consecutive policies and hence the diversity of the policies in the window, possibly making past samples falling outside the clipping range more frequently.
We conjecture that the primary cause is the severely reduced amount of fresh data available to the critic. Since the critic is updated exclusively on the most recent dataset $\cD_k$, reducing new interactions to $H/\omega$ may directly hinder it, possibly leading to degraded advantage estimates that in turn corrupt actor updates. We anticipate that this is not mitigated by fitting the critic on the whole $\cD(\omega_{k})$ (see Appendix~\ref{apx:exp-on-off}).