3D-LLM
AdvancedA large language model that takes 3D scene features directly as input to answer spatial questions and break down tasks.
3D-LLM was proposed in July 2023 by Yining Hong, Chuang Gan, and colleagues from UCLA, UMass Amherst, MIT, and other institutions, a NeurIPS 2023 spotlight. LLMs and vision-language models only handle text or 2D images, and struggle with inherently three-dimensional concepts like spatial relationships and layout. 3D-LLM renders a 3D scene from multiple viewpoints into images, extracts features with a 2D feature extractor, and maps them back onto 3D points to get semantically rich 3D features, which then feed into an off-the-shelf vision-language model such as BLIP-2 for training; it adds a 3D localization mechanism to help the model understand position. The authors designed three prompting methods and collected more than 300,000 3D-language data points covering scene captioning, 3D question answering, task decomposition, localization, and navigation. It beat the previous best method's BLEU-1 score on ScanQA by 9%. It's one of the earlier works connecting 3D scenes to large models, and later work like 3D-VLA and LEO continues this direction.
ExampleGiven a 3D scan of a room, 3D-LLM can answer “which side of the sofa is the fridge on,” or break down “make breakfast” into steps like walking to the kitchen, opening the fridge, and taking out eggs.
- Also called
- 3D-LLM: Injecting the 3D World into Large Language Models
- Related
- 3D VLA · Multimodal Large Language Model · Spatial Reasoning · 3D Visual Grounding · LEO (BIGAI) · ScanNet
- Sources
- 3D-LLM: Injecting the 3D World into Large Language Models (arXiv 2307.12981)
3D-LLM 代码仓库 (GitHub) (Chinese) - As of
- 2023-12