Interactive World Model
可交互世界模型AdvancedA world model that takes an action as input at every step and generates the next frame from it, usable as a neural simulator.
An ordinary video generation model generates a whole clip at once from a piece of text, with no way to intervene partway through; an interactive world model instead takes an action at every step — a keypress, a robot control command, or a text event — and generates the following frame from it, effectively acting as a simulator implemented with a neural network. 2023's UniSim combines many kinds of data to learn how the world visually responds to human and robot actions, and trains policies inside it that deploy zero-shot to the real world; Google DeepMind's 2024 Genie (11 billion parameters) trains only on unlabeled internet video with no action labels, achieving frame-by-frame control through a latent action model (which automatically infers 'what action was taken' from consecutive frames); Genie 3, released in August 2025, can interact in real time at 720p and 24 frames per second while keeping the scene consistent for several minutes. For embodied AI, it can be used to evaluate policies, generate training data, or let an agent learn inside the generated world.
ExampleDeepMind placed its SIMA agent inside a world generated by Genie 3: the agent issues navigation actions like moving forward or turning, Genie 3 generates the corresponding view in real time, and the agent uses it to work toward a given goal.
- Also called
- Action-conditioned Video Model, Generative Interactive Environment
- Related
- World Model · Genie 3 · UniSim · Latent Action Model · Video Generation Model · World-Model-based Policy Evaluation
- Sources
- Genie: Generative Interactive Environments (arXiv:2402.15391)
Learning Interactive Real-World Simulators (UniSim, arXiv:2310.06114)
Genie 3: A new frontier for world models(Google DeepMind 博客) (Chinese) - As of
- 2025-08