OK-Robot
AdvancedA 2024 NYU pick-and-place system built by combining off-the-shelf vision-language models with navigation and grasping modules.
OK-Robot was proposed by Lerrel Pinto's group at NYU in collaboration with Meta (Peiqi Liu, Mahi Shafiullah, and others), posted to arXiv in January 2024, published at RSS 2024. “OK” stands for Open Knowledge, meaning models pretrained openly on internet-scale data. It trains no new end-to-end policy; instead, it combines off-the-shelf modules: first, a lidar-equipped iPhone (using the Record3D app) scans the room once to build a voxel map carrying semantic features like CLIP's; on receiving an instruction like “put this object somewhere,” it finds the object in the map, navigates to it, computes a grasp pose with AnyGrasp, then navigates to the target location and places it down. The whole system runs on a Hello Robot Stretch, with no retraining needed when moving to a new home. The paper's focus is summarizing which details actually determine success or failure when assembling a system like this, and it tallies the causes of failure.
ExampleAcross 171 pick-and-place attempts in 10 real homes in New York, overall success rate was 58.5%, rising to 82% in tidier environments; the most common failures were the semantic map locating the wrong object (9.3%), a difficult grasp pose (8.0%), and hardware issues (7.5%).
- Also called
- What Really Matters in Integrating Open-Knowledge Models for Robotics
- Related
- Open-vocabulary · Zero-shot · Mobile Manipulation · AnyGrasp · Hello Robot Stretch · Semantic Map
- Sources
- OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics (arXiv 2401.12202)
OK-Robot 项目页 (Chinese)
ok-robot/ok-robot GitHub 仓库 (Chinese) - As of
- 2024-07