Embodied AI Glossary中文

Visual Dexterity

Visual Dexterity(视觉手内重定向)Advanced

MIT's system that uses a single depth camera to let a low-cost dexterous hand reorient unseen objects in the air in real time.

Visual Dexterity is work by Tao Chen, Pulkit Agrawal, and colleagues at MIT CSAIL's Improbable AI Lab, posted to arXiv in November 2022 and published in Science Robotics in 2023. In-hand reorientation means turning an object to any target orientation using only the fingers, with no tabletop to rest on — a hard problem because of complex contact and self-occlusion by the hand. The method first trains a 'teacher' in simulation with reinforcement learning that can read the object's true state, then distills it into a 'student' that works only from depth-camera point clouds (processed with a sparse convolutional network, running at about 12 Hz in real time). Training used about 150 objects, on hardware built from the open-source D'Claw three-fingered (9-DoF) or four-fingered hand, for a total system cost under $5,000; it can even reorient objects while palm-down, fighting gravity, with a median completion time of about 7 seconds. This shows that simulation training combined with visual input can generalize to novel object shapes, making it a landmark in in-hand manipulation after Dactyl.

ExampleWith its palm facing down, the robot hand holds a plastic duck it never saw during training in midair and reorients it to a target pose using only depth-camera point clouds; the duck was dropped in 56% of trials, and when it wasn't dropped, about 75% of trials landed within 23 degrees of the target.

Also called
Visual Dexterity: In-Hand Reorientation of Novel and Complex Object Shapes
Related
In-hand Manipulation · Dexterous Manipulation · Teacher-Student Distillation · Sim-to-Real Transfer · Dactyl · HORA
Sources
Visual Dexterity: In-Hand Reorientation of Novel and Complex Object Shapes (arXiv 2211.11744)
Visual Dexterity 项目主页 (Chinese)
As of
2023-11

See it in the full glossary →