RoboPoint
AdvancedA vision-language model that points to where to place or grasp something on an image, following a language instruction.
RoboPoint was released by the University of Washington, NVIDIA, the Allen Institute for AI, and others in June 2024, accepted at CoRL 2024. It focuses on spatial affordance: given an image and an instruction (such as 'put it in the empty spot to the right of the plate'), it marks a set of feasible target locations on the image as 2D points, which are then projected into 3D using a depth map and handed to a motion planner for execution. Its training data is entirely synthesized automatically from procedurally generated 3D scenes, requiring no real-robot data or human demonstrations; the language backbone is Vicuna-13B. The paper reports 21.8% higher spatial-affordance prediction accuracy than GPT-4o and the visual-prompting method PIVOT, and 30.5% higher downstream task success, with applications in manipulation, navigation, and AR assistance. This 'output a point, then hand off to a planner' pattern is a common intermediate representation between VLMs and robots.
ExampleFor the instruction 'put the cup in the empty spot between the two books,' RoboPoint outputs a set of points on the image falling within the empty spot, and the robot arm uses one of them as the placement target.
- Also called
- RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
- Related
- Affordance · Pointing · Spatial Reasoning · Intermediate Representation · Synthetic Data · PIVOT
- Sources
- arXiv 2406.10721: RoboPoint
RoboPoint 项目主页 (Chinese) - As of
- 2024-11