Embodied AI Glossary中文

UniPi

Advanced

A method that first generates a video of a task being completed from text, then infers the robot's actions from that video.

UniPi was released by MIT, Google Brain, UC Berkeley, and the University of Alberta (Yilun Du, Pieter Abbeel, and others) in January 2023, published at NeurIPS 2023. It reframes decision-making as 'text-conditioned video generation': given a text goal and the current image, a video diffusion model generates a future video of the task being completed, which serves as the plan, and an inverse dynamics model (which computes the action between two adjacent frames) then derives the robot action the robot should take from each pair of consecutive frames. Because different robots and environments are all unified into images, the model can share knowledge across tasks; text goals can be freely combined for compositional generalization; and pretraining on internet image-text-video data also improves generalization to new instructions. UniPi is an early representative of the 'video generation model as policy' approach, and later work such as UniSim and Video Prediction Policy follows a similar idea.

ExampleGiven 'put the red block in the blue bowl,' UniPi first generates a short video of a robot completing this task, then an inverse dynamics model derives the actions frame by frame from that video for execution.

Also called
Universal Policies via Text-Guided Video Generation, UniPi: Learning Universal Policies via Text-Guided Video Generation
Related
UniSim · Video Generation Model · Inverse Dynamics Model · Video Prediction Policy · SuSIE · World Model
Sources
Learning Universal Policies via Text-Guided Video Generation (arXiv 2302.00111)
UniPi project page
As of
2023-11

See it in the full glossary →