Gen2Act
AdvancedA manipulation method that first generates a video of a human performing the task, then has the robot execute it by following that video.
Gen2Act was proposed in September 2024 by researchers at Google DeepMind, Carnegie Mellon University, and Stanford. It splits “follow a language instruction to manipulate an object” into two steps: first, a video-generation model trained on web video generates, zero-shot with no fine-tuning, a video of a human completing the task in the current scene; then a robot policy watches this video and executes it, with a point-trajectory prediction loss added during training (predicting how points in the image move) so the policy learns to read motion information out of the video. Generalizing across objects and actions is handed off to the generative model, which has seen a vast amount of web video, while the robot data only has to teach “translating the action in the video into a robot action” — which lets it handle object categories and actions that never appear in the robot data itself.
ExampleAction types absent from the robot data, such as stirring in a circle or dragging an object in a new direction, can be carried out with the help of a generated human video; chaining several such steps together can also accomplish long-horizon tasks like “clear the table” or “make coffee.”
- Also called
- Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation
- Related
- Video Generation Model · Human Video Data · Tracking Any Point · Behavior Cloning · UniPi · DreamGen
- Sources
- Gen2Act (arXiv:2409.16283)
Gen2Act 项目主页 (Chinese) - As of
- 2024-09