Online Meta-Critic Learning for Off-Policy Actor-Critic Methods (2003.05334v2)

Published 11 Mar 2020 in cs.LG and stat.ML

Abstract: Off-Policy Actor-Critic (Off-PAC) methods have proven successful in a variety of continuous control tasks. Normally, the critic's action-value function is updated using temporal-difference, and the critic in turn provides a loss for the actor that trains it to take actions with higher expected return. In this paper, we introduce a novel and flexible meta-critic that observes the learning process and meta-learns an additional loss for the actor that accelerates and improves actor-critic learning. Compared to the vanilla critic, the meta-critic network is explicitly trained to accelerate the learning process; and compared to existing meta-learning algorithms, meta-critic is rapidly learned online for a single task, rather than slowly over a family of tasks. Crucially, our meta-critic framework is designed for off-policy based learners, which currently provide state-of-the-art reinforcement learning sample efficiency. We demonstrate that online meta-critic learning leads to improvements in avariety of continuous control environments when combined with contemporary Off-PAC methods DDPG, TD3 and the state-of-the-art SAC.

Authors (5)

Wei Zhou (311 papers)
Yiying Li (12 papers)
Yongxin Yang (73 papers)
Huaimin Wang (37 papers)
Timothy M. Hospedales (69 papers)

Citations (39)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Related Papers

Off-Policy Actor-Critic (2012)
Deep Exploration with PAC-Bayes (2024)
OPAC: Opportunistic Actor-Critic (2020)
Doubly Robust Off-Policy Actor-Critic Algorithms for Reinforcement Learning (2019)
The Actor-Advisor: Policy Gradient With Off-Policy Advice (2019)