2000 character limit reached
Imitation Learning by Reinforcement Learning
Published 10 Aug 2021 in stat.ML and cs.LG | (2108.04763v2)
Abstract: Imitation learning algorithms learn a policy from demonstrations of expert behavior. We show that, for deterministic experts, imitation learning can be done by reduction to reinforcement learning with a stationary reward. Our theoretical analysis both certifies the recovery of expert reward and bounds the total variation distance between the expert and the imitation learner, showing a link to adversarial imitation learning. We conduct experiments which confirm that our reduction works well in practice for continuous control tasks.
Paper Prompts
Sign up for free to create and run prompts on this paper.