Embodied AI Glossary中文

GR-RL

字节 GR-RLAdvanced

ByteDance Seed uses reinforcement learning to turn a generalist VLA into a specialist, the first learned policy to lace a shoe autonomously.

GR-RL was released by ByteDance's Seed team in December 2025, starting from the generalist model GR-3. Its premise is that ordinary VLA training assumes human demonstrations are optimal, but in fine-grained, long-horizon dexterous tasks, demonstrations often contain jitter and unnecessary motion. GR-RL works in three steps: first, offline reinforcement learning with sparse rewards, using the learned Q-value as an estimate of “task progress” to cut out segments that don't contribute to progress; then morphological symmetry augmentation — flipping images left-right, swapping the left and right wrist cameras, and rewriting left/right words in the instruction to match — to expand the data; and finally online reinforcement learning, which trains a latent-space noise predictor that steers a flow-matching policy's output toward higher reward, aligning behavior between training and actual deployment. The experimental platform is ByteMini-v2, a wheeled robot with two 7-degree-of-freedom arms.

ExampleShoe-lacing: threading a lace through a shoe's eyelets in sequence requires long-horizon planning, millimeter-level precision, and compliant handling of a soft lace. GR-RL reaches an 83.3% success rate, which the authors say is the first learned policy able to complete this task autonomously.

Also called
Going Dexterous and Precise for Long-Horizon Robotic Manipulation
Related
Seed GR-3 · Reinforcement Fine-Tuning (RL Fine-Tuning) · Offline Reinforcement Learning · Noise-Space Policy Steering · Data Curation · Long-horizon Task
Sources
GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation (arXiv 2512.01801)
GR-RL 论文 HTML 全文 (Chinese)
As of
2025-12

See it in the full glossary →