Interference from Cross-Task Experience Replay in Lifelong RL
Determine whether replaying samples from previously encountered tasks interferes with the online gradient updates when learning a new task in lifelong reinforcement learning settings that employ experience replay (e.g., reservoir sampling).
References
We conjecture that replaying samples from other tasks can impose some interference on the online updates when learning a new task.
— A Dirichlet Process Mixture of Robust Task Models for Scalable Lifelong Reinforcement Learning
(2205.10787 - Wang et al., 2022) in Section 4.1 (Simple 2D Navigation)
This was beneficial for overall stability, but it also degraded performance in the harder exploration games such as Pitfall. We were not able to find an easy remedy for this so we did not investigate further in this direction.
— Human-level Atari 200x faster
(2209.07550 - Kapturowski et al., 2022) in Appendix, Section 'Other Things we Tried', subsection 'Mixture of Online and Replay Data'