GR-1 (ByteDance)
字节 GR-1GR-1AdvancedA late-2023 ByteDance manipulation model that first learns to predict future video frames, then learns to output actions.
GR-1 was released by ByteDance Research in December 2023, published at ICLR 2024. It is a GPT-style Transformer: it takes in a language instruction, a history of images, and robot state, and outputs both a future frame and a robot action at the same time. Training happens in two steps: first, video-prediction pretraining on large-scale video, learning only “what will the next frame look like,” which needs no action labels; then fine-tuning on robot data, so the model outputs actions while still predicting frames. The authors' reasoning is that predicting video teaches the model how objects get pushed and how they change, knowledge that is useful for manipulation. On the CALVIN benchmark, it raised the success rate from 88.9% to 94.9%, and zero-shot generalization success on unseen scenes from 53.3% to 85.4%. It is the first generation of ByteDance's GR series, followed by GR-2 and GR-3. Note that this is not the same GR-1 as Fourier's humanoid robot.
ExampleIn the CALVIN simulated kitchen tabletop, GR-1 follows language instructions to complete a sequence of sub-tasks in a row, such as “open the drawer,” “push the blue block to the left,” and “turn on the light bulb.”
- Also called
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation
- Related
- Video Prediction Model · GR-2 (ByteDance) · Seed GR-3 · CALVIN Benchmark · Language-conditioned Policy · Pre-training
- Sources
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation (arXiv 2312.13139)
GR-1 项目页 (Chinese)
bytedance/GR-1 GitHub 仓库(ICLR 2024) (Chinese) - As of
- 2024-01