Reinforcement Fine-Tuning (RL Fine-Tuning)
强化学习微调RFTCommonFurther improving a pretrained or supervised-fine-tuned model with reinforcement learning driven by a reward signal.
Reinforcement fine-tuning takes a model that's already been pretrained or supervised-fine-tuned (SFT) as its starting policy, lets it generate its own answers or actions, scores them with a reward, and updates the parameters with an algorithm such as policy gradient. Unlike SFT, which only imitates demonstrations, this lets the model learn from its own successes and failures, discovering behaviors the demonstrations never showed. The term “RFT” became popular through OpenAI's reinforcement fine-tuning service for its o-series reasoning models: the user supplies a grader, and the model samples several answers per question and reinforces the higher-scoring ones. Since 2025, a wave of work has applied this to VLA models, usually with a binary success/failure reward and algorithms like PPO or GRPO; because a flow-matching action head can't compute an action's probability directly, it needs special adaptation, as in πRL.
ExampleπRL fine-tunes flow-matching VLAs with online reinforcement learning: π0 first reaches 57.6% success on LIBERO after SFT on a handful of demonstrations, then rises to 97.6% after reinforcement learning; π0.5 goes from 77.1% to 98.3%. SimpleVLA-RL instead trains OpenVLA-OFT with just a “1 for success, 0 for failure” reward combined with GRPO.
- Also called
- RFT, RL Fine-tuning, RL Post-training
- Related
- Post-training · Supervised Fine-Tuning · Proximal Policy Optimization · Group Relative Policy Optimization · πRL · SimpleVLA-RL
- Sources
- OpenAI API Docs: Reinforcement fine-tuning
πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning - As of
- 2026-01