Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
80 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Learning 3D Semantic Segmentation with only 2D Image Supervision (2110.11325v1)

Published 21 Oct 2021 in cs.CV

Abstract: With the recent growth of urban mapping and autonomous driving efforts, there has been an explosion of raw 3D data collected from terrestrial platforms with lidar scanners and color cameras. However, due to high labeling costs, ground-truth 3D semantic segmentation annotations are limited in both quantity and geographic diversity, while also being difficult to transfer across sensors. In contrast, large image collections with ground-truth semantic segmentations are readily available for diverse sets of scenes. In this paper, we investigate how to use only those labeled 2D image collections to supervise training 3D semantic segmentation models. Our approach is to train a 3D model from pseudo-labels derived from 2D semantic image segmentations using multiview fusion. We address several novel issues with this approach, including how to select trusted pseudo-labels, how to sample 3D scenes with rare object categories, and how to decouple input features from 2D images from pseudo-labels during training. The proposed network architecture, 2D3DNet, achieves significantly better performance (+6.2-11.4 mIoU) than baselines during experiments on a new urban dataset with lidar and images captured in 20 cities across 5 continents.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (9)
  1. Kyle Genova (21 papers)
  2. Xiaoqi Yin (8 papers)
  3. Abhijit Kundu (16 papers)
  4. Caroline Pantofaru (15 papers)
  5. Forrester Cole (27 papers)
  6. Avneesh Sud (16 papers)
  7. Brian Brewington (2 papers)
  8. Brian Shucker (1 paper)
  9. Thomas Funkhouser (66 papers)
Citations (66)