Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
80 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning (2311.18651v1)

Published 30 Nov 2023 in cs.CV

Abstract: Recent advances in Large Multimodal Models (LMM) have made it possible for various applications in human-machine interactions. However, developing LMMs that can comprehend, reason, and plan in complex and diverse 3D environments remains a challenging topic, especially considering the demand for understanding permutation-invariant point cloud 3D representations of the 3D scene. Existing works seek help from multi-view images, and project 2D features to 3D space as 3D scene representations. This, however, leads to huge computational overhead and performance degradation. In this paper, we present LL3DA, a Large Language 3D Assistant that takes point cloud as direct input and respond to both textual-instructions and visual-prompts. This help LMMs better comprehend human interactions and further help to remove the ambiguities in cluttered 3D scenes. Experiments show that LL3DA achieves remarkable results, and surpasses various 3D vision-LLMs on both 3D Dense Captioning and 3D Question Answering.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (9)
  1. Sijin Chen (12 papers)
  2. Xin Chen (456 papers)
  3. Chi Zhang (566 papers)
  4. Mingsheng Li (9 papers)
  5. Gang Yu (114 papers)
  6. Hao Fei (105 papers)
  7. Hongyuan Zhu (36 papers)
  8. Jiayuan Fan (29 papers)
  9. Tao Chen (397 papers)
Citations (42)
X Twitter Logo Streamline Icon: https://streamlinehq.com
Youtube Logo Streamline Icon: https://streamlinehq.com