veRL (Volcano Engine Reinforcement Learning)
veRLAdvancedByteDance's open-source large-model reinforcement learning framework, also used for VLA RL fine-tuning.
veRL is a reinforcement learning training library open-sourced by ByteDance's Seed team (released under the Volcano Engine name), and is the open-source implementation of the HybridFlow paper (EuroSys 2025). Doing RL on a large model means repeatedly alternating between two things: using the current model to generate a batch of samples (rollout), then using those samples to update the parameters — and the two steps want different kinds of parallelism. veRL hands generation to inference engines such as vLLM or SGLang, hands training to FSDP or Megatron, and handles the weight synchronization and scheduling between the two, with PPO and GRPO built in. It originally targeted language models, and was later adopted by the embodied-AI community for online reinforcement-learning fine-tuning of VLAs.
ExampleSimpleVLA-RL is built on veRL: it runs parallel rollouts in simulators such as LIBERO, using task success or failure as the reward to fine-tune OpenVLA-OFT with GRPO-style reinforcement learning.
- Also called
- verl, HybridFlow
- Related
- Reinforcement Fine-Tuning (RL Fine-Tuning) · Group Relative Policy Optimization · Proximal Policy Optimization · vLLM · SimpleVLA-RL · Ray
- Sources
- volcengine/verl GitHub
HybridFlow: A Flexible and Efficient RLHF Framework (arXiv)