F3RM
AdvancedA 2023 MIT method that distills 2D features like CLIP into a 3D field, letting a robot grasp objects from language instructions with few demonstrations.
F3RM (Feature Fields for Robotic Manipulation) comes from William Shen, Ge Yang, Phillip Isola, and colleagues at MIT CSAIL, published at CoRL 2023. 2D image models like CLIP and DINO understand semantics but have no notion of an object's precise 3D position and shape, and robotic grasping can't do without geometry. F3RM first takes multiple photos of a scene and trains a neural radiance field (NeRF, a method that reconstructs a 3D scene from multi-view photos), while simultaneously distilling features from models like CLIP into that same 3D field, producing a distilled feature field where every point in space carries a semantic feature. A robot then optimizes its gripper pose within this field, learning 6-DOF grasping and placing from just a handful of demonstrations; the target object can be specified with free-form text, and the approach generalizes to objects it has never seen. It shares its core idea with LERF (which embeds language features into a radiance field), and is a representative example of carrying 2D foundation-model knowledge into 3D for manipulation.
ExampleGiven just two demonstrations per task, such as grasping a mug by its handle or its rim, the robot can pick up mugs of different colors and sizes; typing a short text description lets it pick out the matching object in a scene to grasp.
- Also called
- Feature Fields for Robotic Manipulation, Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation
- Related
- Distilled Feature Fields · Neural Radiance Fields · CLIP · LERF · Open-vocabulary · Few-shot
- Sources
- Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation (arXiv 2308.07931)
F3RM 项目主页 (Chinese) - As of
- 2023-11