Embodied AI Glossary中文

RIPT-VLA

RIPT-VLA(交互式后训练)Advanced

A method that post-trains a pretrained VLA with reinforcement learning using only a binary success/failure reward.

RIPT-VLA was released by Shuhan Tan, Philipp Krähenbühl, and colleagues at UT Austin in May 2025. VLAs are normally pretrained and then supervised-fine-tuned on expert demonstrations, but this works poorly when demonstrations are scarce. RIPT-VLA adds a third stage, 'interactive post-training': the model repeatedly attempts the task in the environment and is trained with reinforcement learning using only a binary reward for task success or failure. It samples the same initial state multiple times and estimates the advantage using the average score of the other attempts in the same group as a baseline (a leave-one-out estimator), then updates the policy with PPO; groups where every attempt succeeds or every attempt fails carry no learning signal and are discarded and resampled. It improves the QueST model by 21.2% and pushes the 7B OpenVLA-OFT to 97.5% on LIBERO; given just a single demonstration, it can take a model with 4% success up to 97% within 15 iterations.

ExampleA VLA fine-tuned on just one demonstration starts at only 4% success; letting it repeatedly attempt the task in simulation and scoring each attempt as success or failure raises its success rate to 97% after 15 iterations.

Also called
Reinforcement Interactive Post-Training, RIPT-VLA: Interactive Post-Training for Vision-Language-Action Models
Related
Reinforcement Fine-Tuning (RL Fine-Tuning) · Post-training · OpenVLA-OFT · LIBERO Benchmark · Proximal Policy Optimization · SimpleVLA-RL
Sources
Interactive Post-Training for Vision-Language-Action Models (arXiv 2505.17016)
As of
2025-05

See it in the full glossary →