Seed GR-3
字节 GR-3GR-3CommonByteDance Seed's 2025 general-purpose robot VLA, about 4 billion parameters, controlling its own ByteMini bimanual robot.
GR-3 is a vision-language-action model the ByteDance Seed team released in July 2025, following GR-1 and GR-2. It has about 4 billion parameters, using Qwen2.5-VL-3B as its vision-language backbone, followed by a diffusion Transformer action head trained with flow matching that generates one action chunk at a time. Training draws on three kinds of data: robot trajectories for imitation learning; web image-text data co-trained in to preserve understanding of new objects and abstract instructions; and human trajectories captured with a VR device for few-shot adaptation — the paper reports that adding just 10 demonstrations per unseen object raised pick-and-place success from 57.8% to 86.7%. Its companion robot, ByteMini, is a 22-degree-of-freedom bimanual mobile platform, and GR-3 outperforms π0 on long-horizon tasks such as clearing a table and hanging up clothes.
ExampleIn the clothes-hanging task, GR-3 coordinates ByteMini's two arms to thread a hanger into a garment and then hang it on a rack — an example of bimanual deformable-object manipulation.
- Also called
- ByteDance Seed GR-3, Generalist Robot Model 3
- Related
- GR-2 (ByteDance) · GR-1 (ByteDance) · ByteDance Seed · Vision-Language-Action Model · Flow Matching · Co-training
- Sources
- GR-3 Technical Report (arXiv:2507.15493)
- As of
- 2025-07