When to Go, and When to Explore: The Benefit of Post-Exploration in Intrinsic Motivation (2203.16311v2)

Published 29 Mar 2022 in cs.LG and cs.AI

Abstract: Go-Explore achieved breakthrough performance on challenging reinforcement learning (RL) tasks with sparse rewards. The key insight of Go-Explore was that successful exploration requires an agent to first return to an interesting state ('Go'), and only then explore into unknown terrain ('Explore'). We refer to such exploration after a goal is reached as 'post-exploration'. In this paper we present a systematic study of post-exploration, answering open questions that the Go-Explore paper did not answer yet. First, we study the isolated potential of post-exploration, by turning it on and off within the same algorithm. Subsequently, we introduce new methodology to adaptively decide when to post-explore and for how long to post-explore. Experiments on a range of MiniGrid environments show that post-exploration indeed boosts performance (with a bigger impact than tuning regular exploration parameters), and this effect is further enhanced by adaptively deciding when and for how long to post-explore. In short, our work identifies adaptive post-exploration as a promising direction for RL exploration research.

Authors (4)

Zhao Yang (75 papers)
Thomas M. Moerland (24 papers)
Mike Preuss (39 papers)
Aske Plaat (76 papers)

Citations (1)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

When to Go, and When to Explore: The Benefit of Post-Exploration in Intrinsic Motivation (2203.16311v2)

Summary

Related Papers