UniSim
AdvancedA learned 'real-world simulator,' built from a video generation model, that responds to actions.
UniSim was released by UC Berkeley, Google DeepMind, and MIT (Sherry Yang, Yilun Du, Pieter Abbeel, and others) in October 2023, winning the ICLR 2024 Outstanding Paper Award. It uses a video generation model to learn a real-world simulator: given the current image and an action, it generates the resulting image. The action can either be a high-level language instruction like 'open the drawer' or a low-level control command like 'move to a given coordinate.' No single dataset covers all of this information, so the authors combined image data (rich in object variety), robot data (rich in actions), and navigation data (rich in movement), among others, for training. Once trained, vision-language policies and reinforcement learning policies can be trained inside UniSim and then transferred zero-shot to real robots; the data it generates can also help train other models, such as video captioning. UniSim is an early landmark for interactive world models and neural simulators.
ExampleGiven a kitchen image and the instruction 'open the drawer,' UniSim generates the subsequent video of the drawer being pulled open; a reinforcement learning policy can then practice through trial and error inside these generated images.
- Also called
- Universal Simulator, Learning Interactive Real-World Simulators
- Related
- UniPi · Interactive World Model · Neural Simulator · World Model · Genie (Original) · Video Generation Model
- Sources
- Learning Interactive Real-World Simulators (arXiv 2310.06114)
UniSim project page - As of
- 2024-05