LERF
AdvancedEmbeds CLIP language features into a NeRF, so objects in a 3D scene can be found with natural language.
LERF was proposed by Justin Kerr, Chung Min Kim, Angjoo Kanazawa, and colleagues at UC Berkeley, published at ICCV 2023 as an oral presentation. On top of a neural radiance field (NeRF, which uses a network to represent color and density at every point in a scene), it additionally learns a multi-scale language field: a point in space maps to a CLIP vector at each of several scales, supervised during training by CLIP features extracted from an image pyramid across multiple viewpoints, and regularized with DINO features to sharpen object boundaries. Once built, feeding in any text produces a relevance heatmap rendered in 3D space, with no detection box, segmentation mask, or model fine-tuning needed. It is an early landmark for bringing a vision-language model’s open-vocabulary ability into 3D scenes, with code integrated into Nerfstudio.
ExampleLERF-TOGO (2023) first uses LERF to localize an object part like “the handle of the cup” in a scene, then ranks the candidate grasps from an off-the-shelf grasp planner; across 31 real objects, it picked the correct part 81% of the time, with a 69% grasp success rate.
- Also called
- Language Embedded Radiance Fields, LERF-TOGO
- Related
- Neural Radiance Fields · CLIP · Distilled Feature Fields · F3RM · Open-vocabulary · 3D Visual Grounding
- Sources
- LERF: Language Embedded Radiance Fields (ICCV 2023 项目主页) (Chinese)
Language Embedded Radiance Fields for Zero-Shot Task-Oriented Grasping (LERF-TOGO, arXiv 2309.07970)