Embodied AI Glossary中文

SimpleVLA-RL

Advanced

An open-source framework that runs large-scale online reinforcement learning on a VLA using only a success/failure reward.

SimpleVLA-RL was released and open-sourced in September 2025 by Tsinghua University, the Shanghai Artificial Intelligence Laboratory, and other institutions, accepted at ICLR 2026. VLAs are usually trained with supervised fine-tuning (SFT) to imitate human demonstrations, but high-quality real-robot demonstrations are expensive, and policies trained this way tend to fail when they hit out-of-distribution situations. SimpleVLA-RL borrows the idea of using reinforcement learning to improve reasoning in large language models, adapting veRL, an RL framework built for large models, for VLAs — adding interactive trajectory sampling, parallel rendering across many environments, and distributed training; the reward looks only at whether the task ultimately succeeded (0 or 1), with no hand-designed reward shaping. Starting from OpenVLA-OFT as the base model, it reaches state-of-the-art results on LIBERO at the time and beats π0 on RoboTwin 1.0 and 2.0; when SFT is done with just one demonstration per task, LIBERO-Long success can go from 17.3% to 91.7%. The authors also observed a 'pushcut' phenomenon, where RL training discovers behaviors absent from the demonstrations, such as pushing an object into place instead of picking it up as demonstrated.

ExampleIn LIBERO simulation, OpenVLA-OFT is first fine-tuned on one demonstration per task until it occasionally succeeds, then lets it repeatedly try in a large number of parallel environments, rewarded only by final success or failure, which sharply raises its success rate.

Also called
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
Related
Reinforcement Fine-Tuning (RL Fine-Tuning) · OpenVLA-OFT · veRL (Volcano Engine Reinforcement Learning) · Reinforcement Learning with Verifiable Rewards · RoboTwin · LIBERO Benchmark
Sources
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning (arXiv 2509.09674)
PRIME-RL/SimpleVLA-RL (GitHub)
As of
2026-01

See it in the full glossary →