Embodied AI Glossary中文

Seed GR-3

字节 GR-3GR-3Common

ByteDance Seed's 2025 general-purpose robot VLA, about 4 billion parameters, controlling its own ByteMini bimanual robot.

GR-3 is a vision-language-action model the ByteDance Seed team released in July 2025, following GR-1 and GR-2. It has about 4 billion parameters, using Qwen2.5-VL-3B as its vision-language backbone, followed by a diffusion Transformer action head trained with flow matching that generates one action chunk at a time. Training draws on three kinds of data: robot trajectories for imitation learning; web image-text data co-trained in to preserve understanding of new objects and abstract instructions; and human trajectories captured with a VR device for few-shot adaptation — the paper reports that adding just 10 demonstrations per unseen object raised pick-and-place success from 57.8% to 86.7%. Its companion robot, ByteMini, is a 22-degree-of-freedom bimanual mobile platform, and GR-3 outperforms π0 on long-horizon tasks such as clearing a table and hanging up clothes.

ExampleIn the clothes-hanging task, GR-3 coordinates ByteMini's two arms to thread a hanger into a garment and then hang it on a rack — an example of bimanual deformable-object manipulation.

Also called
ByteDance Seed GR-3, Generalist Robot Model 3
Related
GR-2 (ByteDance) · GR-1 (ByteDance) · ByteDance Seed · Vision-Language-Action Model · Flow Matching · Co-training
Sources
GR-3 Technical Report (arXiv:2507.15493)
As of
2025-07

See it in the full glossary →