Embodied AI Glossary中文

Gato

Advanced

A 2022 DeepMind generalist agent whose single set of weights can play games, caption images, chat, and control a robot arm.

Gato is a generalist agent DeepMind released in May 2022, in a paper titled A Generalist Agent. The mainstream approach at the time was one model per task; Gato instead used a single roughly 1.2-billion-parameter Transformer, with one set of weights, to cover 604 tasks: playing Atari games, captioning images, holding conversations, controlling various robots in simulation, and stacking blocks with a real robot arm. The key idea is converting every modality into tokens: text is split by a tokenizer, images are cut into patches, and discrete or continuous values like joint angles and button presses are also encoded as tokens, all concatenated into a single sequence and predicted one at a time like a language model, with the training loss computed only on the text and action tokens. At inference, the model autoregressively samples action tokens one at a time, with a context window of 1,024 tokens. It is an early representative of the “one large model controlling many embodiments” idea, and DeepMind's later RoboCat carried over Gato's architecture.

ExampleThe same Gato model presses controller buttons in an Atari game one moment, generates a caption for a photo the next, and then controls a real robot arm to stack colored blocks.

Also called
A Generalist Agent
Related
Generalist Policy · Transformer · Token · RoboCat · Multi-Task Learning · Cross-Embodiment
Sources
A Generalist Agent (arXiv 2205.06175)
A Generalist Agent(Google DeepMind 博客) (Chinese)
As of
2022-05

See it in the full glossary →