ConRFT
AdvancedA method that fine-tunes a VLA with reinforcement learning using a consistency policy, offline first, then online on the real robot.
ConRFT was proposed in February 2025 by Dongbin Zhao's group at the Institute of Automation, Chinese Academy of Sciences, published at RSS 2025. Supervised fine-tuning of a VLA on a small number of demonstrations often isn't reliable enough for contact-rich real-robot tasks, while running reinforcement learning directly on the real robot is slow and unsafe. ConRFT works in two stages. The offline stage (Cal-ConRFT) combines behavior cloning with Q-learning to learn a reasonably stable policy and value estimate from a small number of demonstrations. The online stage (HIL-ConRFT) continues fine-tuning with reinforcement learning on the real robot, letting a person take over and correct it at any time to keep exploration safe. The action head uses a consistency policy (a diffusion-style model that can generate an action in one or a few steps), which fits well with reinforcement learning. Built on Octo-small, after 45–90 minutes of online fine-tuning across 8 real tasks, average success rate reached 96.3%, 144% higher than pure supervised fine-tuning.
ExampleTasks included picking up a banana, opening a drawer, putting bread in a toaster, mounting a car wheel, and hanging a Chinese knot; during online training, the operator could take over whenever the robot was about to make a mistake.
- Also called
- Cal-ConRFT, HIL-ConRFT
- Related
- Reinforcement Fine-Tuning (RL Fine-Tuning) · Consistency Policy · HIL-SERL · Human-in-the-Loop · Octo · Real-World Reinforcement Learning
- Sources
- ConRFT (arXiv 2502.05450)
ConRFT 项目主页 (Chinese) - As of
- 2025-04