Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Random Expert Distillation: Imitation Learning via Expert Policy Support Estimation (1905.06750v2)

Published 16 May 2019 in cs.LG and stat.ML

Abstract: We consider the problem of imitation learning from a finite set of expert trajectories, without access to reinforcement signals. The classical approach of extracting the expert's reward function via inverse reinforcement learning, followed by reinforcement learning is indirect and may be computationally expensive. Recent generative adversarial methods based on matching the policy distribution between the expert and the agent could be unstable during training. We propose a new framework for imitation learning by estimating the support of the expert policy to compute a fixed reward function, which allows us to re-frame imitation learning within the standard reinforcement learning setting. We demonstrate the efficacy of our reward function on both discrete and continuous domains, achieving comparable or better performance than the state of the art under different reinforcement learning algorithms.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Ruohan Wang (22 papers)
  2. Carlo Ciliberto (41 papers)
  3. Pierluigi Amadori (4 papers)
  4. Yiannis Demiris (26 papers)
Citations (61)

Summary

We haven't generated a summary for this paper yet.