Papers
Topics
Authors
Recent
Search
2000 character limit reached

ODTFormer: Efficient Obstacle Detection and Tracking with Stereo Cameras Based on Transformer

Published 21 Mar 2024 in cs.RO and cs.CV | (2403.14626v3)

Abstract: Obstacle detection and tracking represent a critical component in robot autonomous navigation. In this paper, we propose ODTFormer, a Transformer-based model to address both obstacle detection and tracking problems. For the detection task, our approach leverages deformable attention to construct a 3D cost volume, which is decoded progressively in the form of voxel occupancy grids. We further track the obstacles by matching the voxels between consecutive frames. The entire model can be optimized in an end-to-end manner. Through extensive experiments on DrivingStereo and KITTI benchmarks, our model achieves state-of-the-art performance in the obstacle detection task. We also report comparable accuracy to state-of-the-art obstacle tracking models while requiring only a fraction of their computation cost, typically ten-fold to twenty-fold less. The code and model weights will be publicly released.

Authors (3)
Definition Search Book Streamline Icon: https://streamlinehq.com
References (42)
  1. J. Frey et al., “Locomotion Policy Guided Traversability Learning using Volumetric Representations of Complex Environments,” Aug. 2022, arXiv:2203.15854 [cs].
  2. Y. F. Chen et al., “Socially aware motion planning with deep reinforcement learning,” in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Sep. 2017, pp. 1343–1350.
  3. N. U. Akmandor et al., “Deep Reinforcement Learning based Robot Navigation in Dynamic Environments using Occupancy Values of Motion Primitives,” in IROS, 2022.
  4. L. Zhao et al., “EE\mathrm{E}roman_E(2)-equivariant graph planning for navigation,” IEEE Robotics and Automation Letters, vol. 9, no. 4, 2024.
  5. H. Li et al., “Stereonavnet: Learning to navigate using stereo cameras with auxiliary occupancy voxels,” Mar. 2024, arXiv:2403.12039 [cs].
  6. Z. Li et al., “BEVFormer: Learning Bird’s-Eye-View Representation from Multi-camera Images via Spatiotemporal Transformers,” in Computer Vision – ECCV 2022, S. Avidan et al., Eds., 2022, pp. 1–18.
  7. Y. Li et al., “VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene Completion,” 2023, pp. 9087–9098.
  8. Y. Wei et al., “SurroundOcc: Multi-camera 3D Occupancy Prediction for Autonomous Driving,” 2023, pp. 21 729–21 740.
  9. Y. Huang et al., “Tri-Perspective View for Vision-Based 3D Semantic Occupancy Prediction,” 2023, pp. 9223–9232.
  10. T. Eppenberger et al., “Leveraging Stereo-Camera Data for Real-Time Dynamic Obstacle Detection and Tracking,” in IROS, 2020.
  11. Z. Xu et al., “A real-time dynamic obstacle tracking and mapping system for UAV navigation and collision avoidance with an RGB-D camera,” in ICRA, 2023.
  12. M. Lu et al., “Perception and Avoidance of Multiple Small Fast Moving Objects for Quadrotors With Only Low-Cost RGBD Camera,” IEEE RA-L, vol. 7, no. 4, Oct. 2022.
  13. Y. Wang et al., “Autonomous Flights in Dynamic Environments with Onboard Vision,” in IROS, 2021.
  14. H. Li et al., “Stereovoxelnet: Real-time obstacle detection based on occupancy voxels from a stereo camera using deep neural networks,” in 2023 IEEE International Conference on Robotics and Automation (ICRA), 2023.
  15. S. Kareer et al., “ViNL: Visual Navigation and Locomotion Over Obstacles,” Jan. 2023, arXiv:2210.14791 [cs].
  16. B. Zhou et al., “RAPTOR: Robust and Perception-aware Trajectory Replanning for Quadrotor Fast Flight,” Jul. 2020, arXiv:2007.03465.
  17. A. Loquercio et al., “Learning high-speed flight in the wild,” Science Robotics, vol. 6, no. 59, p. eabg5810, 2021.
  18. D. Falanga et al., “How Fast Is Too Fast? The Role of Perception Latency in High-Speed Sense and Avoid,” IEEE Robotics and Automation Letters, vol. 4, no. 2, pp. 1884–1891, Apr. 2019.
  19. Y. Wang et al., “Pseudo-LiDAR From Visual Depth Estimation: Bridging the Gap in 3D Object Detection for Autonomous Driving,” 2019, pp. 8445–8453.
  20. S. Vedula et al., “Three-dimensional scene flow,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 27, no. 3, pp. 475–480, Mar. 2005.
  21. H. Jiang et al., “SENSE: a Shared Encoder Network for Scene-flow Estimation,” Oct. 2019, arXiv:1910.12361 [cs].
  22. Z. Teed and J. Deng, “RAFT-3D: Scene Flow using Rigid-Motion Embeddings,” Apr. 2021, arXiv:2012.00726 [cs].
  23. A. Vaswani et al., “Attention is All you Need,” in Advances in Neural Information Processing Systems, vol. 30, 2017.
  24. F. Shamsafar et al., “MobileStereoNet: Towards Lightweight Deep Networks for Stereo Matching,” 2022, pp. 2417–2426.
  25. Z. Shen et al., “CFNet: Cascade and Fused Cost Volume for Robust Stereo Matching,” 2021, pp. 13 906–13 915.
  26. B. Liu et al., “Local Similarity Pattern and Cost Self-Reassembling for Deep Stereo Matching Networks,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 2, pp. 1647–1655, Jun. 2022, number: 2.
  27. G. Xu et al., “Attention Concatenation Volume for Accurate and Efficient Stereo Matching,” 2022, pp. 12 981–12 990.
  28. X. Zhu et al., “Deformable detr: Deformable transformers for end-to-end object detection,” arXiv preprint arXiv:2010.04159, 2020.
  29. G. Yang et al., “Drivingstereo: A large-scale dataset for stereo matching in autonomous driving scenarios,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019.
  30. M. Menze and A. Geiger, “Object scene flow for autonomous vehicles,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3061–3070.
  31. M. Tan and Q. Le, “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” in Proceedings of the 36th International Conference on Machine Learning.   PMLR, May 2019, pp. 6105–6114.
  32. T.-Y. Lin et al., “Feature pyramid networks for object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2117–2125.
  33. H. Hirschmuller, “Accurate and efficient stereo processing by semi-global matching and mutual information,” in CVPR, vol. 2, 2005.
  34. R. Saxena et al., “Pwoc-3d: Deep occlusion-aware end-to-end scene flow estimation,” in 2019 IEEE Intelligent Vehicles Symposium (IV).   IEEE, 2019, pp. 324–331.
  35. I. Shepel et al., “Occupancy grid generation with dynamic obstacle segmentation in stereo images,” IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 9, pp. 14 779–14 789, 2021.
  36. H. Wu et al., “Casa: A cascade attention network for 3-d object detection from lidar point clouds,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–11, 2022.
  37. ——, “Virtual sparse convolution for multimodal 3d object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 21 653–21 662.
  38. X. Wang et al., “You only need two detectors to achieve multi-modal 3d multi-object tracking,” arXiv preprint arXiv:2304.08709, 2023.
  39. Y. Xie et al., “Pixel-Aligned Recurrent Queries for Multi-View 3D Object Detection,” in ICCV, 2023.
  40. I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” Sep. 2018.
  41. H. Jiang et al., “Sense: A shared encoder network for scene-flow estimation,” in ICCV, 2019.
  42. N. Mayer et al., “A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation,” in CVPR, 2016.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 3 tweets with 0 likes about this paper.