Genie (Original)
Genie(初代)AdvancedGoogle DeepMind's model that learns a “playable world” from unlabeled video, a generative interactive environment.
Genie was released by Google DeepMind in February 2024, with 11 billion parameters, and is the first generative interactive environment trained in an unsupervised way from nothing but unlabeled internet video. It has three components: a spatiotemporal video tokenizer that compresses video into discrete tokens; a latent action model that infers, from a pair of consecutive frames, which action occurred, out of a small set of discrete action codes (8 in the paper); and an autoregressive dynamics model that predicts the next frame from the current frame and the action. The key point is that no action labels are needed during training, yet a person can still control the generated world frame by frame. It was trained mainly on 2D platform-jumping game videos, and was also validated on RT-1 robot video; the latent-action idea was later carried over by robot-pretraining work such as LAPA.
ExampleGive Genie a hand-drawn sketch, and it will turn it into a 2D platform game you can play frame by frame with your actions.
- Also called
- Genie 1, Generative Interactive Environments
- Related
- Genie 2 · Genie 3 · Latent Action Model · Video Tokenizer · Interactive World Model · LAPA
- Sources
- Genie: Generative Interactive Environments (arXiv:2402.15391)
Genie 项目主页 (Chinese) - As of
- 2024-02