Embodied AI Glossary中文

GR-2 (ByteDance)

字节 GR-2GR-2Advanced

ByteDance's second-generation manipulation model, released in 2024, pretrained first on 38 million web videos before learning to act.

GR-2 was released by ByteDance Research in October 2024, an upgrade over GR-1 that the authors call a “video-language-action model.” The first stage does video-generation pretraining on 38 million internet videos (more than 50 billion tokens), drawn from everyday human-activity video sources like HowTo100M, Ego4D, Something-Something V2, and EPIC-KITCHENS, teaching the model how footage will unfold next; the second stage fine-tunes on robot trajectories, predicting future frames and an action trajectory at the same time. Images are cut into discrete tokens with VQGAN, and the action trajectory is generated by a conditional variational autoencoder (CVAE). The paper reports an average 97.7% success rate across more than 100 tasks, along with strong generalization to new backgrounds, environments, objects, and tasks. It was deployed on a real Kinova Gen3 arm, with trajectory optimization and a real-time-tracking whole-body control algorithm carrying out the actions.

ExampleIn an industrial bin-picking experiment requiring picking a specified item out of a cluttered bin, across 122 object types (67 of them unseen during training), GR-2 reached an average 79.0% success rate, versus just 33.3% for GR-1.

Also called
A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Related
GR-1 (ByteDance) · Seed GR-3 · Video Generation Model · Pretraining on Human Videos · Conditional Variational Autoencoder · Bin Picking
Sources
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation (arXiv 2410.06158)
GR-2 论文 HTML 全文 (Chinese)
As of
2024-10

See it in the full glossary →