SeaBird: Segmentation in Bird's View with Dice Loss Improves Monocular 3D Detection of Large Objects
Abstract: Monocular 3D detectors achieve remarkable performance on cars and smaller objects. However, their performance drops on larger objects, leading to fatal accidents. Some attribute the failures to training data scarcity or their receptive field requirements of large objects. In this paper, we highlight this understudied problem of generalization to large objects. We find that modern frontal detectors struggle to generalize to large objects even on nearly balanced datasets. We argue that the cause of failure is the sensitivity of depth regression losses to noise of larger objects. To bridge this gap, we comprehensively investigate regression and dice losses, examining their robustness under varying error levels and object sizes. We mathematically prove that the dice loss leads to superior noise-robustness and model convergence for large objects compared to regression losses for a simplified case. Leveraging our theoretical insights, we propose SeaBird (Segmentation in Bird's View) as the first step towards generalizing to large objects. SeaBird effectively integrates BEV segmentation on foreground objects for 3D detection, with the segmentation head trained with the dice loss. SeaBird achieves SoTA results on the KITTI-360 leaderboard and improves existing detectors on the nuScenes leaderboard, particularly for large objects. Code and models at https://github.com/abhi1kumar/SeaBird
- Augmented reality meets computer vision: Efficient data generation for urban driving scenes. IJCV, 2018.
- SemanticKITTI: A dataset for semantic scene understanding of LiDAR sequences. In ICCV, 2019.
- Zygmunt Birnbaum. An inequality for Mill’s ratio. The Annals of Mathematical Statistics, 1942.
- M3333D-RPN: Monocular 3333D region proposal network for object detection. In ICCV, 2019.
- Kinematic 3333D object detection in monocular video. In ECCV, 2020.
- Omni3D: A large benchmark and model for 3333D object detection in the wild. In CVPR, 2023.
- nuScenes: A multimodal dataset for autonomous driving. In CVPR, 2020.
- Brittany Caldwell. 2 die when tesla crashes into parked tractor-trailer in florida. https://www.wftv.com/news/local/2-die-when-tesla-crashes-into-parked-tractor-trailer-florida/KJGMHHYTQZA2HNAHWL2OFSVIPM/, 2022. Accessed: 2023-11-06.
- End-to-end object detection with transformers. In ECCV, 2020.
- Deep MANTA: A coarse-to-fine many-task network for joint 2222D and 3333D vehicle analysis from monocular image. In CVPR, 2017.
- Viewpoint equivariance for multi-view 3333D object detection. In CVPR, 2023.
- Monocular 3333D object detection for autonomous driving. In CVPR, 2016.
- DSGN: Deep stereo geometry network for 3333D object detection. In CVPR, 2020a.
- MonoPair: Monocular 3333D object detection using pairwise spatial relationships. In CVPR, 2020b.
- NEAT: Neural attention fields for end-to-end autonomous driving. In ICCV, 2021.
- Depth-discriminative metric learning for monocular 3333D object detection. In NeurIPS, 2023.
- OA-BEV: Bringing object awareness to bird’s-eye-view representation for multi-camera 3333D object detection. arXiv preprint arXiv:2301.05711, 2023.
- MMDetection3D Contributors. MMDetection3D: OpenMMLab next-generation platform for general 3333D object detection. https://github.com/open-mmlab/mmdetection3d, 2020.
- Michael Crawshaw. Multi-task learning with deep neural networks: A survey. arXiv preprint arXiv:2009.09796, 2020.
- SpatialDETR: Robust scalable transformer-based 3333D object detection from multi-view camera images with global cross-sensor attention. In ECCV, 2022.
- Benchmarking robustness of 3333D object detection to common corruptions. In CVPR, 2023.
- Fully sparse 3333D object detection. In NeurIPS, 2022.
- AEDet: Azimuth-invariant multi-view 3333D object detection. arXiv preprint arXiv:2211.12501, 2022.
- Roshan Fernandez. A tesla driver was killed after smashing into a firetruck on a california highway. https://www.npr.org/2023/02/20/1158367204/tesla-driver-killed-california-firetruck-nhtsa, 2023. Accessed: 2023-11-06.
- Are we ready for autonomous driving? the KITTI vision benchmark suite. In CVPR, 2012.
- Ross Girshick. Fast R-CNN. In ICCV, 2015.
- Bird’s-eye-view panoptic segmentation using monocular frontal view images. RAL, 2022.
- Simple-BEV: What really matters for multi-sensor BEV perception? In CoRL, 2022.
- Deep residual learning for image recognition. In CVPR, 2016.
- FIERY: future instance prediction in bird’s-eye view from surround monocular cameras. In ICCV, 2021.
- BEVDet4D: Exploit temporal cues in multi-camera 3333D object detection. arXiv preprint arXiv:2203.17054, 2022.
- BEVDet: High-performance multi-camera 3333D object detection in bird-eye-view. arXiv preprint arXiv:2112.11790, 2021.
- MonoDTR: Monocular 3333D object detection with depth-aware transformer. In CVPR, 2022.
- STXD: Structural and temporal cross-modal distillation for multi-view 3333D object detection. In NeurIPS, 2023.
- MonoUNI: A unified vehicle and infrastructure-side monocular 3333D object detection network with sufficient depth clues. In NeurIPS, 2023.
- Polarformer: Multi-camera 3333D object detection with polar transformers. In AAAI, 2023.
- Predict to Detect: Prediction-guided 3333D object detection using sequential images. In ICCV, 2023.
- Adam: A method for stochastic optimization. In ICLR, 2015.
- Towards viewpoint robustness in Bird’s Eye View segmentation. In ICCV, 2023.
- X3KD: Knowledge distillation across modalities, tasks and stages for multi-camera 3333D object detection. In CVPR, 2023.
- LUVLi face alignment: Estimating landmarks’ location, uncertainty, and visibility likelihood. In CVPR, 2020.
- GrooMeD-NMS: Grouped mathematically differentiable NMS for monocular 3333D object detection. In CVPR, 2021.
- DEVIANT: Depth Equivariant Network for monocular 3333D object detection. In ECCV, 2022.
- A simpler approach to obtaining an 𝒪(1/t)𝒪1𝑡\mathcal{O}\left(1/t\right)caligraphic_O ( 1 / italic_t ) convergence rate for the projected stochastic subgradient method. arXiv preprint arXiv:1212.2002, 2012.
- BAAM: Monocular 3333D pose and shape reconstruction with bi-contextual attention module and attention-guided modeling. In CVPR, 2023.
- Unifying voxel-based representation with transformer for 3333D object detection. In NeurIPS, 2022a.
- BEVStereo: Enhancing depth estimation in multi-view 3333D object detection with dynamic temporal stereo. In AAAI, 2023a.
- BEVDepth: Acquisition of reliable depth for multi-view 3333D object detection. In AAAI, 2023b.
- Fast-BEV: A fast and strong bird’s-eye view perception baseline. In NeurIPS Workshops, 2023c.
- BEVFormer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers. In ECCV, 2022b.
- FB-BEV: BEV representation from forward-backward view transformations. In ICCV, 2023d.
- KITTI-360: A novel dataset and benchmarks for urban scene understanding in 2222D and 3333D. TPAMI, 2022.
- Feature pyramid networks for object detection. In CVPR, 2017.
- Voxel-based 3333D detection and reconstruction of multiple objects from a single image. In NeurIPS, 2021.
- SparseBEV: High-performance sparse 3333D object detection from multi-camera videos. In ICCV, 2023a.
- Monocular 3333D object detection with bounding box denoising in 3333D by perceiver. In ICCV, 2023b.
- PETR: Position embedding transformation for multi-view 3333D object detection. In ECCV, 2022.
- PETRv2: A unified framework for 3333D perception from multi-camera images. In ICCV, 2023c.
- Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021a.
- AutoShape: Real-time shape-aware monocular 3333D object detection. In ICCV, 2021b.
- RADIANT: RADar Image Association Network for 3333D object detection. In AAAI, 2023.
- Decoupled weight decay regularization. In ICLR, 2019.
- Geometry uncertainty projection network for monocular 3333D object detection. In ICCV, 2021.
- DETR4D: Direct multi-view 3333D object detection with sparse attention. arXiv preprint arXiv:2212.07849, 2022.
- Accurate monocular 3333D object detection via color-embedded 3333D reconstruction for autonomous driving. In ICCV, 2019.
- Delving into localization errors for monocular 3333D object detection. In CVPR, 2021.
- 3333D object detection from images for autonomous driving: A survey. TPAMI, 2023a.
- Towards fair and comprehensive comparisons for image-based 3333D object detection. In ICCV, 2023b.
- Vision-centric BEV perception: A survey. arXiv preprint arXiv:2208.02797, 2022.
- Symmetry and uncertainty-aware object SLAM for 6666DoF object pose estimation. In CVPR, 2022.
- NeurOCS: Neural NOCS supervision for monocular 3333D object localization. In CVPR, 2023.
- Rotation matters: Generalized monocular 3333D object detection for various camera systems. arXiv preprint arXiv:2310.05366, 2023.
- Cross-view semantic segmentation for sensing surroundings. RAL, 2020.
- Is Pseudo-LiDAR needed for monocular 3333D object detection? In ICCV, 2021.
- Time will tell: New outlooks and a baseline for temporal multi-view 3333D object detection. In ICLR, 2023.
- Pix2Pose: Pixel-wise coordinate regression of objects for 6666D pose estimation. In ICCV, 2019.
- PyTorch: An imperative style, high-performance deep learning library. In NeurIPS, 2019.
- From contours to 3333D object detection and pose estimation. In ICCV, 2011.
- Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3333D. In ECCV, 2020.
- Deep hough voting for 3333D object detection in point clouds. In ICCV, 2019.
- Categorical depth distribution network for monocular 3333D object detection. In CVPR, 2021.
- Predicting semantic map representations from images using pyramid occupancy networks. In CVPR, 2020.
- Translating images into maps. In ICRA, 2022.
- Robotic grasping of novel objects using vision. IJRR, 2008.
- Pegasos: Primal estimated sub-gradient solver for SVM. In ICML, 2007.
- PointRCNN: 3333D object proposal generation and detection from point cloud. In CVPR, 2019.
- Distance-normalized unified representation for monocular 3333D object detection. In ECCV, 2020.
- Multivariate probabilistic monocular 3333D object detection. In WACV, 2023.
- 3DPPE: 3333D point positional encoding for multi-camera 3333D object detection transformers. In ICCV, 2023.
- Scalability in perception for autonomous driving: Waymo open dataset. In CVPR, 2020.
- EfficientDet: Scalable and efficient object detection. In CVPR, 2020.
- StreamPETR: Exploring object-centric temporal modeling for efficient multi-view 3333D object detection. In ICCV, 2023a.
- FCOS3D: Fully convolutional one-stage monocular 3333D object detection. In ICCV Workshops, 2021a.
- Probabilistic and geometric depth: Detecting objects in perspective. In CoRL, 2021b.
- Segmentation can aid detection: Segmentation-guided single stage detection for 3333D point cloud. Electronics, 2023b.
- Pseudo-LiDAR from visual depth estimation: Bridging the gap in 3333D object detection for autonomous driving. In CVPR, 2019.
- DETR3D: 3333D object detection from multi-view images via 3333D-to-2222D queries. In CoRL, 2021c.
- FrustumFormer: Adaptive instance-aware resampling for multi-view 3333D detection. In CVPR, 2023c.
- Object as Query: Lifting any 2222D object detector to 3333D detection. In ICCV, 2023d.
- DistillBEV: Boosting multi-camera 3333D object detection with cross-modal knowledge distillation. In ICCV, 2023e.
- STS: Surround-view temporal stereo for multi-view 3333D detection. In AAAI, 2023f.
- Chen Wu. Waymo keynote talk, CVPR workshop on autonomous driving at 17:20. https://www.youtube.com/watch?v=fXsbI2VkHgc, 2023. Accessed: 2023-11-11.
- M2̂BEV: Multi-camera joint 3333D detection and segmentation with unified birds-eye view representation. arXiv preprint arXiv:2204.05088, 2022.
- CAPE: Camera view position embedding for multi-view 3333D object detection. In CVPR, 2023.
- MonoNeRD: NeRF-like representations for monocular 3333D object detection. In ICCV, 2023.
- LiDAR-based 3333D object detection via hybrid 2222D semantic scene generation. arXiv preprint arXiv:2304.01519, 2023a.
- Parametric depth based feature representation learning for object detection and segmentation in bird’s-eye view. In ICCV, 2023b.
- Oriented object detection in aerial images with box boundary-aware vectors. In WACV, 2021.
- Center-based 3333D object detection and tracking. In CVPR, 2021.
- PoseCNN: A convolutional neural network for 6666D object pose estimation in cluttered scenes. In RSS, 2018.
- Learning enriched features for fast image restoration and enhancement. TPAMI, 2022.
- DA-BEV: Depth aware BEV transformer for 3333D object detection. arXiv preprint arXiv:2302.13002, 2023a.
- SA-BEV: Generating semantic-aware bird’s-eye-view feature for multi-view 3333D object detection. In ICCV, 2023b.
- MonoDETR: Depth-guided transformer for monocular 3333D object detection. In ICCV, 2023c.
- Objects are different: Flexible monocular 3333D object detection. In CVPR, 2021.
- BEVerse: Unified perception and prediction in birds-eye-view for vision-centric autonomous driving. arXiv preprint arXiv:2205.09743, 2022.
- Cross-view transformers for real-time map-view semantic segmentation. In CVPR, 2022.
- MonoEF: Extrinsic parameter free monocular 3333D object detection. TPAMI, 2021.
- Class-balanced grouping and sampling for point cloud 3333D object detection. In CVPR Workshop, 2019.
- Understanding the robustness of 3333D object detection with bird’s-eye-view representations in autonomous driving. In CVPR, 2023.
- Temporal enhanced training of multi-view 3333D object detector via historical object prediction. In ICCV, 2023.
Paper Prompts
Sign up for free to create and run prompts on this paper.