Embodied AI Glossary中文

3D Visual Grounding

3D视觉定位Advanced

Finds the object in a 3D scene that a sentence describes.

3D visual grounding brings 2D visual grounding — using text to draw a box around a target in an image — into 3D scenes: given a scene’s point cloud or multi-view images plus a description such as “the trash can next to the chair by the window,” it outputs the target object’s 3D bounding box. The Chinese term for this, 视觉定位, is also commonly used for a camera estimating its own position, but here it means grounding. The main difficulty is that a scene often has multiple objects of the same category, so the model must understand spatial relationships like “on the left,” “the biggest one,” or “next to the door” to tell them apart. ScanRefer and ReferIt3D, both published at ECCV 2020, built benchmarks on ScanNet indoor scans: ScanRefer has about 52,000 descriptions covering 11,000 objects, while ReferIt3D provides two sets, Nr3D (natural descriptions) and Sr3D (spatial-relation descriptions). For robots, this is the key step that grounds a language instruction in an actual object.

ExampleA user says “hand me the blue cup on the far right of the desk”; the robot runs 3D visual grounding on the reconstructed room point cloud to get that cup’s 3D box, then plans a grasp.

Also called
3D Grounding, 3D Referring Expression Grounding
Related
Visual Grounding · ScanNet · 3D Object Detection · 3D Scene Graph · Open-Vocabulary Object Detection · Language Grounding
Sources
ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language (ECCV 2020)
ReferIt3D: Neural Listeners for Fine-Grained 3D Object Identification in Real-World Scenes (ECCV 2020)

See it in the full glossary →