Embodied AI Glossary中文

MimicPlay

Advanced

An imitation-learning method that learns high-level plans from video of human hands playing freely, then low-level actions from a little teleoperation data.

MimicPlay was released in February 2023 by Chen Wang, Linxi Fan, Fei-Fei Li, Yuke Zhu, and colleagues at Stanford, NVIDIA, and other institutions, an oral presentation at CoRL 2023. Collecting data purely through teleoperation is expensive for long-horizon tasks (tasks requiring many steps done in sequence). MimicPlay uses a hierarchical structure: the high level learns a latent plan from “human play data” — video of a person freely moving objects around a scene by hand — representing what the 3D trajectory of a human hand should look like given a goal image; the low level trains a visuomotor policy on only a small number of teleoperation demonstrations, outputting robot actions that follow this plan. Human-hand video is cheap and works across embodiments, filling exactly the gap left by scarce robot data. Across 14 real long-horizon manipulation tasks, it beat the contemporary baselines on success rate, generalization, and robustness to disturbance.

ExampleFor a multi-step task like “open the microwave, put the bowl inside, then close the door,” the high level predicts the 3D trajectory a human hand should follow given the goal image, and the low-level policy drives the robot arm step by step along that trajectory to complete the task.

Also called
Long-Horizon Imitation Learning by Watching Human Play
Related
Imitation Learning · Long-horizon Task · Play Data · Human Video Data · Hierarchical Architecture · Visuomotor Policy
Sources
MimicPlay: Long-Horizon Imitation Learning by Watching Human Play (arXiv 2302.12422)
MimicPlay project page
As of
2023-10

See it in the full glossary →