Reinforcement Learning with Verifiable Rewards
基于可验证奖励的强化学习RLVRAdvancedTraining a model with reinforcement learning using rewards a program can automatically check as correct or wrong.
The name RLVR comes from the Allen Institute for AI's (Ai2) Tülu 3 paper in November 2024: for tasks such as math problems or automatically checkable instruction constraints, a rule or program judges whether the model's output is correct, gives a fixed reward for a correct answer and zero otherwise, and then updates the model with an algorithm such as PPO. It needs no separately trained reward model, the network that scores answers in RLHF, so the reward source is more reliable and harder for the model to game. DeepSeek-R1 (2025) used a rule-based accuracy reward to train reasoning ability, bringing this approach wide attention. Applied to robots, “did the task succeed or not” is a naturally verifiable reward, and work such as SimpleVLA-RL and VLA-RFT uses a binary success/failure reward to fine-tune VLAs with reinforcement learning.
ExampleSimpleVLA-RL fine-tunes OpenVLA-OFT with reinforcement learning using only an outcome reward — success is rewarded, failure is not — reaching leading results on LIBERO, beating π0 on RoboTwin 1.0 and 2.0, and also outperforming a purely supervised-fine-tuned version on a real robot.
- Also called
- RLVR
- Related
- Reinforcement Fine-Tuning (RL Fine-Tuning) · Reward Model · Reinforcement Learning from Human Feedback · Group Relative Policy Optimization · Sparse Reward · SimpleVLA-RL
- Sources
- Lambert et al. 2024: Tülu 3: Pushing Frontiers in Open Language Model Post-Training
DeepSeek-AI 2025: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Li et al. 2025: SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning - As of
- 2025-09