Papers
Topics
Authors
Recent
Search
2000 character limit reached

Scaling Motion Forecasting Models with Ensemble Distillation

Published 5 Apr 2024 in cs.RO and cs.LG | (2404.03843v2)

Abstract: Motion forecasting has become an increasingly critical component of autonomous robotic systems. Onboard compute budgets typically limit the accuracy of real-time systems. In this work we propose methods of improving motion forecasting systems subject to limited compute budgets by combining model ensemble and distillation techniques. The use of ensembles of deep neural networks has been shown to improve generalization accuracy in many application domains. We first demonstrate significant performance gains by creating a large ensemble of optimized single models. We then develop a generalized framework to distill motion forecasting model ensembles into small student models which retain high performance with a fraction of the computing cost. For this study we focus on the task of motion forecasting using real world data from autonomous driving systems. We develop ensemble models that are very competitive on the Waymo Open Motion Dataset (WOMD) and Argoverse leaderboards. From these ensembles, we train distilled student models which have high performance at a fraction of the compute costs. These experiments demonstrate distillation from ensembles as an effective method for improving accuracy of predictive models for robotic systems with limited compute budgets.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (43)
  1. Social lstm: Human trajectory prediction in crowded spaces. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 961–971, 2016.
  2. Towards understanding ensemble, knowledge distillation and self-distillation in deep learning, 2020.
  3. Towards understanding ensemble, knowledge distillation and self-distillation in deep learning. arXiv preprint, (2012.09816), December 2020.
  4. L. Breiman. Stacked regressions. Machine Learning, 24:49–64, 2004.
  5. Spagnn: Spatially-aware graph neural networks for relational behavior forecasting from sensor data. In IEEE Intl. Conf. on Robotics and Automation. IEEE, 2020.
  6. Intentnet: Learning to predict intention from raw sensor data. CoRR, abs/2101.07907, 2021.
  7. Multipath: Multiple probabilistic anchor trajectory hypotheses for behavior prediction. In CoRL, 2019.
  8. Argoverse: 3d tracking and forecasting with rich maps. CoRR, abs/1911.02620, 2019.
  9. Multimodal trajectory predictions for autonomous driving using deep convolutional networks. In 2019 International Conference on Robotics and Automation (ICRA), pages 2090–2096. IEEE, 2019.
  10. Thomas G. Dietterich. Ensemble methods in machine learning. In Multiple Classifier Systems, 2000.
  11. Large scale interactive motion forecasting for autonomous driving : The waymo open motion dataset. CoRR, abs/2104.10133, 2021.
  12. Deep ensembles: A loss landscape perspective. ArXiv, abs/1912.02757, 2019.
  13. Born again neural networks. In ICML, 2018.
  14. VectorNet: Encoding hd maps and agent dynamics from vectorized representation. In CVPR, 2020.
  15. Densetnt: End-to-end trajectory prediction from dense goal sets. CoRR, abs/2108.09640, 2021.
  16. Social gan: Socially acceptable trajectories with generative adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018.
  17. Neural network ensembles. IEEE Trans. Pattern Anal. Mach. Intell., 12:993–1001, 1990.
  18. Distilling the knowledge in a neural network. ArXiv, abs/1503.02531, 2015.
  19. Rules of the road: Predicting driving behavior with a convolutional model of semantic interactions. CoRR, abs/1906.08945, 2019.
  20. One thousand and one hours: Self-driving motion prediction dataset. CoRR, abs/2006.14480, 2020.
  21. Scaling laws for neural language models. ArXiv, abs/2001.08361, 2020.
  22. What-if motion prediction for autonomous driving. ArXiv, 2020.
  23. Sequence-level knowledge distillation. In EMNLP, 2016.
  24. Neural network ensembles, cross validation, and active learning. In NIPS, 1994.
  25. DESIRE: distant future prediction in dynamic scenes with interacting agents. CoRR, abs/1704.04394, 2017.
  26. Learning small-size dnn with output-distribution-based criteria. In INTERSPEECH, 2014.
  27. Learning lane graph representations for motion forecasting. arXiv preprint arXiv:2007.13732, 2020.
  28. Decoupled weight decay regularization. In ICLR, 2019.
  29. Popular ensemble methods: An empirical study. J. Artif. Intell. Res., 11:169–198, 1999.
  30. Ensemble distribution distillation, 2019.
  31. Multi-head attention for multi-modal joint vehicle motion forecasting. 2020 IEEE International Conference on Robotics and Automation (ICRA), pages 9638–9644, 2020.
  32. Wayformer: Motion forecasting via simple & efficient attention networks. ArXiv, abs/2207.05844, 2022.
  33. Precog: Prediction conditioned on goals in visual multi-agent settings. In ECCV, 2019.
  34. Trajectron++: Multi-agent generative trajectory forecasting with heterogeneous data for control. CoRR, abs/2001.03093, 2020.
  35. MultiPATH: Multiple probabilistic anchor trajectory hypotheses for behavior prediction. In Conf. on Robot Learning, 2019.
  36. Boosting the margin: A new explanation for the effectiveness of voting methods. In ICML, 1997.
  37. Motion transformer with global intention localization and local movement refinement. Advances in Neural Information Processing Systems, 2022.
  38. Narrowing the coordinate-frame gap in behavior prediction models: Distillation for efficient and accurate scene-centric motion forecasting. 2022 International Conference on Robotics and Automation (ICRA), pages 653–659, 2022.
  39. Multipath++: Efficient information fusion and trajectory aggregation for behavior prediction. In ICRA, 2021.
  40. Attention is all you need. In NeurIPS, 2017.
  41. Argoverse 2: Next generation datasets for self-driving perception and forecasting. arXiv preprint arXiv:2301.00493, 2023.
  42. Tnt: Target-driven trajectory prediction. arXiv preprint arXiv:2008.08294, 2020.
  43. Understanding knowledge distillation in non-autoregressive machine translation. ArXiv, abs/1911.02727, 2020.
Citations (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.