Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
102 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Refer-it-in-RGBD: A Bottom-up Approach for 3D Visual Grounding in RGBD Images (2103.07894v3)

Published 14 Mar 2021 in cs.CV

Abstract: Grounding referring expressions in RGBD image has been an emerging field. We present a novel task of 3D visual grounding in single-view RGBD image where the referred objects are often only partially scanned due to occlusion. In contrast to previous works that directly generate object proposals for grounding in the 3D scenes, we propose a bottom-up approach to gradually aggregate context-aware information, effectively addressing the challenge posed by the partial geometry. Our approach first fuses the language and the visual features at the bottom level to generate a heatmap that coarsely localizes the relevant regions in the RGBD image. Then our approach conducts an adaptive feature learning based on the heatmap and performs the object-level matching with another visio-linguistic fusion to finally ground the referred object. We evaluate the proposed method by comparing to the state-of-the-art methods on both the RGBD images extracted from the ScanRefer dataset and our newly collected SUNRefer dataset. Experiments show that our method outperforms the previous methods by a large margin (by 11.2% and 15.6% [email protected]) on both datasets.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (6)
  1. Haolin Liu (31 papers)
  2. Anran Lin (3 papers)
  3. Xiaoguang Han (118 papers)
  4. Lei Yang (372 papers)
  5. Yizhou Yu (148 papers)
  6. Shuguang Cui (275 papers)
Citations (37)

Summary

We haven't generated a summary for this paper yet.