Embodied AI Glossary中文

VLA-RFT

Advanced

A method that runs reinforcement learning inside a learned world model to make a VLA more robust in just a few hundred steps.

VLA-RFT was proposed in October 2025 by Westlake University, Zhejiang University, the OpenHelix team, and others. A VLA trained with imitation learning alone tends to accumulate errors and fail under perturbations; reinforcement learning can help, but real-robot interaction is expensive and traditional simulators still have a sim-to-real gap. The method first trains a roughly 138-million-parameter autoregressive world model on real interaction data, predicting future frames from the current image and action, and uses it as a controllable simulator; the policy then rolls out whole trajectories inside it, with reward computed from the pixel-level (L1) and perceptual (LPIPS) error between the predicted frames and the expert reference trajectory's frames — a verifiable reward — and the policy is updated with GRPO (Group Relative Policy Optimization). The base policy is VLA-Adapter. Fine-tuning for under 400 steps already beats the supervised fine-tuning baseline, and the result is more robust to perturbations in object position and initial state, showing that a world model can serve as a practical environment for VLA post-training.

ExampleOn LIBERO, the VLA-Adapter supervised fine-tuning baseline reaches 86.6% average success; after 400 steps of reinforcement fine-tuning inside the world model, this rises to 91.1%.

Also called
VLA Reinforcement Fine-Tuning in World Simulators, VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
Related
World Model · Reinforcement Fine-Tuning (RL Fine-Tuning) · Group Relative Policy Optimization · VLA-Adapter · WMPO · Learning in Imagination
Sources
VLA-RFT (arXiv 2510.00406)
VLA-RFT 项目主页 (Chinese)
As of
2025-10

See it in the full glossary →