OpenVLA-OFT
OFTCommonStanford's fine-tuning recipe for VLAs that makes OpenVLA generate actions 26 times faster, with a higher success rate.
OpenVLA-OFT is a VLA fine-tuning recipe proposed by Moo Jin Kim, Chelsea Finn, and Percy Liang at Stanford in February 2025. The original OpenVLA discretizes actions into tokens and outputs them one at a time, autoregressively, which is slow and hard to use for high-frequency control. OFT changes four things during fine-tuning: parallel decoding (computing all actions in a single forward pass), action chunking (predicting multiple steps at once), a continuous action representation, and an L1 regression loss; OFT+ additionally adds FiLM feature modulation to strengthen how well the model follows language instructions. The result: average success rate on LIBERO rose from 76.5% to 97.1%, action-generation throughput increased 26x, and on a real ALOHA bimanual robot it beat other fine-tuned VLAs, including π0 and RDT-1B. It's now commonly used as a strong baseline for VLA fine-tuning.
ExampleWhen fine-tuning OpenVLA on LIBERO, switching the output from discrete action tokens emitted one at a time to an entire continuous action chunk regressed in a single forward pass sharply raises inference throughput and also improves success rate.
- Also called
- OFT, OpenVLA-OFT+, Optimized Fine-Tuning (VLA recipe)
- Related
- OpenVLA · Vision-Language-Action Model · Action Chunking · Parallel Decoding · Feature-wise Linear Modulation · LIBERO Benchmark
- Sources
- OpenVLA-OFT 项目主页 (Chinese)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (arXiv:2502.19645) - As of
- 2025-04