Source of diminishing returns from larger reuse windows
Determine whether the diminishing benefits of increasing the sample-reuse window in Proximal Policy Optimization arise from the increasing proportion of stale advantages and samples collected under increasingly distant behavioral policies within each mini-batch.
References
We conjecture this is due to the increasing proportion of stale advantages and of samples collected under increasingly distant policies within each mini-batch, which gradually offsets the benefit of additional data as the window grows.
We conjecture this is due to the larger policy changes occurring during early training: consecutive policies differ the most in this phase, so that, with a large reuse window, older samples carry the most stale advantages and come from the most distant behavioral policies.