Embodied AI Glossary中文

LERF

Advanced

Embeds CLIP language features into a NeRF, so objects in a 3D scene can be found with natural language.

LERF was proposed by Justin Kerr, Chung Min Kim, Angjoo Kanazawa, and colleagues at UC Berkeley, published at ICCV 2023 as an oral presentation. On top of a neural radiance field (NeRF, which uses a network to represent color and density at every point in a scene), it additionally learns a multi-scale language field: a point in space maps to a CLIP vector at each of several scales, supervised during training by CLIP features extracted from an image pyramid across multiple viewpoints, and regularized with DINO features to sharpen object boundaries. Once built, feeding in any text produces a relevance heatmap rendered in 3D space, with no detection box, segmentation mask, or model fine-tuning needed. It is an early landmark for bringing a vision-language model’s open-vocabulary ability into 3D scenes, with code integrated into Nerfstudio.

ExampleLERF-TOGO (2023) first uses LERF to localize an object part like “the handle of the cup” in a scene, then ranks the candidate grasps from an off-the-shelf grasp planner; across 31 real objects, it picked the correct part 81% of the time, with a 69% grasp success rate.

Also called
Language Embedded Radiance Fields, LERF-TOGO
Related
Neural Radiance Fields · CLIP · Distilled Feature Fields · F3RM · Open-vocabulary · 3D Visual Grounding
Sources
LERF: Language Embedded Radiance Fields (ICCV 2023 项目主页) (Chinese)
Language Embedded Radiance Fields for Zero-Shot Task-Oriented Grasping (LERF-TOGO, arXiv 2309.07970)

See it in the full glossary →