Embodied AI Glossary中文

Hand-Eye Coordination

手眼协调Common

Using what a camera sees to guide an arm and gripper's motion in real time, adjusting as it goes.

Hand-eye coordination originally describes a capability of humans and animals: the eyes see a target and the hand reaches out to grasp it accurately, correcting itself as it watches. Applied to robots, it means closing the loop between camera images and arm motion. The traditional approach first performs hand-eye calibration, computing the coordinate transform between the camera and the arm, then converts a detected target position into arm coordinates and executes the motion in one shot — any small calibration error causes a missed grasp. In 2016, Sergey Levine and colleagues at Google used 6 to 14 robot arms over two months to collect more than 800,000 grasp attempts, training a convolutional network to judge directly from a single monocular image whether a given gripper motion would succeed, with no camera calibration needed and with continuous adjustment during the grasp itself — a landmark example of learning hand-eye coordination end to end. Note this is a distinct concept from hand-eye calibration.

ExampleAn object gets bumped out of place mid-grasp; the robot keeps correcting the gripper's position based on the camera feed, rather than reaching to a single coordinate computed in advance.

Related
Hand-Eye Calibration · Visual Servoing · Visuomotor Policy · Closed-loop Control · Google Arm Farm · Grasping
Sources
Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection (arXiv 1603.02199)

See it in the full glossary →