Embodied AI Glossary中文

Group Relative Policy Optimization

组相对策略优化GRPOCommon

A reinforcement-learning algorithm that samples a group of outputs for the same input and uses their relative scores instead of a value network.

GRPO was introduced by the DeepSeek team in the DeepSeekMath paper (February 2024) as a variant of PPO (Proximal Policy Optimization). PPO needs a separately trained value network (a “critic,” roughly as large as the policy itself) to estimate a baseline; GRPO removes it. For the same input, it samples a group of outputs, scores each one, and uses “score minus the group's mean, divided by the group's standard deviation” as each output's advantage, saving a large amount of memory. It also adds the KL divergence (a measure of how far the policy has drifted) from a reference model directly into the loss, to constrain how much the policy changes per update. GRPO became widely adopted after DeepSeek-R1 used it to train reasoning ability. The embodied-AI field has begun applying it to RL fine-tuning of VLAs too: several rollouts are sampled for the same task and given a reward of 0 or 1 depending on success.

ExampleSimpleVLA-RL trains OpenVLA-OFT with GRPO: it samples 8 rollouts per task, scores success as 1 and failure as 0, drops the KL term, and filters out groups that are all successes or all failures, raising the average LIBERO success rate from 91.0% to 99.1%.

Also called
GRPO
Related
Proximal Policy Optimization · Advantage Function · Reinforcement Fine-Tuning (RL Fine-Tuning) · Reinforcement Learning with Verifiable Rewards · KL Regularization · SimpleVLA-RL
Sources
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
As of
2025-09

See it in the full glossary →