Papers
Topics
Authors
Recent
Search
2000 character limit reached

Mimicking Better by Matching the Approximate Action Distribution

Published 16 Jun 2023 in cs.LG | (2306.09805v3)

Abstract: In this paper, we introduce MAAD, a novel, sample-efficient on-policy algorithm for Imitation Learning from Observations. MAAD utilizes a surrogate reward signal, which can be derived from various sources such as adversarial games, trajectory matching objectives, or optimal transport criteria. To compensate for the non-availability of expert actions, we rely on an inverse dynamics model that infers plausible actions distribution given the expert's state-state transitions; we regularize the imitator's policy by aligning it to the inferred action distribution. MAAD leads to significantly improved sample efficiency and stability. We demonstrate its effectiveness in a number of MuJoCo environments, both int the OpenAI Gym and the DeepMind Control Suite. We show that it requires considerable fewer interactions to achieve expert performance, outperforming current state-of-the-art on-policy methods. Remarkably, MAAD often stands out as the sole method capable of attaining expert performance levels, underscoring its simplicity and efficacy.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (44)
  1. Ls-iq: Implicit reward regularization for inverse reinforcement learning. arXiv preprint arXiv:2303.00599, 2023.
  2. Analyzing inverse problems with invertible neural networks. 2019.
  3. A framework for behavioural cloning. In Machine Intelligence 15, 1995.
  4. Bishop, C. M. Mixture density networks. 1994.
  5. Sample-efficient imitation learning via generative adversarial nets. In International Conference on Artificial Intelligence and Statistics, 2018.
  6. Openai gym. arXiv preprint arXiv:1606.01540, 2016.
  7. Imitation learning from pixel observations for continuous control, 2022.
  8. The frontier of simulation-based inference. Proceedings of the National Academy of Sciences of the United States of America, 117(48):30055–30062, 2020. ISSN 10916490. doi: 10.1073/pnas.1912789117.
  9. Cuturi, M. Sinkhorn distances: Lightspeed computation of optimal transport. In Burges, C., Bottou, L., Welling, M., Ghahramani, Z., and Weinberger, K. (eds.), Advances in Neural Information Processing Systems, volume 26. Curran Associates, Inc., 2013. URL https://proceedings.neurips.cc/paper_files/paper/2013/file/af21d0c97db2e27e13572cbf59eb343d-Paper.pdf.
  10. Primal wasserstein imitation learning. ICLR, 2021.
  11. Model-based imitation learning by probabilistic trajectory matching. 2013 IEEE International Conference on Robotics and Automation, pp.  1922–1927, 2013.
  12. Cross-domain imitation learning via optimal transport. ArXiv, abs/2110.03684, 2021.
  13. Diagnosing bottlenecks in deep q-learning algorithms. In International Conference on Machine Learning, 2019.
  14. A minimalist approach to offline reinforcement learning. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  15. Off-policy deep reinforcement learning without exploration. In International Conference on Machine Learning, 2018a.
  16. Addressing function approximation error in actor-critic methods. In International Conference on Machine Learning, 2018b.
  17. Automatic posterior transformation for likelihood-free inference. In International Conference on Machine Learning, pp. 2404–2414. PMLR, 2019.
  18. Hybrid reinforcement learning with expert state sequences. In AAAI Conference on Artificial Intelligence, 2019.
  19. Watch and match: Supercharging imitation with regularized optimal transport. In Conference on Robot Learning, 2022.
  20. Stochastic grounded action transformation for robot learning in simulation. 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp.  6106–6111, 2017.
  21. Generative adversarial imitation learning. In Advances in Neural Information Processing Systems (NeurIPS), 2016.
  22. Augmenting gail with bc for sample efficient imitation learning. In Conference on Robot Learning, 2020.
  23. Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning. ArXiv, abs/1809.02925, 2018.
  24. Imitation learning via off-policy distribution matching. In International Conference on Learning Representations, ICLR, 2020.
  25. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. ArXiv, abs/2005.01643, 2020.
  26. Flexible statistical inference for mechanistic models of neural dynamics. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pp.  1289–1299, 2017.
  27. Combining self-supervised learning and imitation for vision-based rope manipulation. 2017 IEEE International Conference on Robotics and Automation (ICRA), pp.  2146–2153, 2017.
  28. Training language models to follow instructions with human feedback. ArXiv, abs/2203.02155, 2022.
  29. Imitation learning with sinkhorn distances. In ECML/PKDD, 2020.
  30. Fast ε𝜀\varepsilonitalic_ε-free inference of simulation models with bayesian conditional density estimation. Advances in neural information processing systems, 29, 2016.
  31. Ridm: Reinforced inverse dynamics modeling for learning from a single observed demonstration. IEEE Robotics and Automation Letters, 5:6262–6269, 2019.
  32. Deepmimic: Example-guided deep reinforcement learning of physics-based character skills. ACM Trans. Graph., 37:143, 2018.
  33. Gromov-wasserstein averaging of kernel and distance matrices. In International Conference on Machine Learning, 2016.
  34. Pomerleau, D. Efficient training of artificial neural networks for autonomous navigation. Neural Comput., 3(1):88–97, 1991.
  35. A reduction of imitation learning and structured prediction to no-regret online learning. In International Conference on Artificial Intelligence and Statistics, 2010.
  36. Proximal policy optimization algorithms. ArXiv, abs/1707.06347, 2017.
  37. Mastering the game of go without human knowledge. Nature, 550:354–359, 2017.
  38. Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, pp.  5026–5033. IEEE, 2012. doi: 10.1109/IROS.2012.6386109.
  39. Behavioral cloning from observation. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI, pp.  4950–4957, 2018.
  40. Generative adversarial imitation from observation. ICML Workshop on Imitation, Intent, and Interaction, arXiv:abs/1807.06158, 2019.
  41. dm_control: Software and tasks for continuous control. Software Impacts, 6:100022, 2020.
  42. Imitation learning from observations by minimizing inverse dynamics disagreement. In Neural Information Processing Systems, 2019.
  43. Planning for sample efficient imitation learning. In Neural Information Processing Systems, 2022.
  44. Off-policy imitation learning from observations. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
Citations (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.