Embodied AI Glossary中文

OpenVLA

Essential

A 7-billion-parameter vision-language-action model that Stanford, UC Berkeley, and collaborators open-sourced in 2024, code and weights included.

OpenVLA was released in June 2024 by researchers from Stanford, UC Berkeley, Toyota Research Institute, Google DeepMind, Physical Intelligence, and MIT, and published at CoRL 2024. It's built on the Prismatic vision-language model: the vision encoder fuses features from SigLIP and DINOv2, the language model is Llama 2 7B, and the total parameter count is about 7 billion. For actions, it follows RT-2's approach of discretizing each action dimension into 256 bins and having the language model predict them as tokens, one at a time. It was trained on 970,000 robot trajectories from the Open X-Embodiment dataset, and beat the 55-billion-parameter closed-source RT-2-X by 16.5 absolute percentage points in success rate across 29 tasks. Just as important, its code and weights are fully open, and it can be fine-tuned on consumer GPUs with LoRA (low-rank adaptation, which trains only a small number of added parameters) — which made it a common starting point and baseline for later VLA research, such as OpenVLA-OFT.

ExampleTo adapt OpenVLA to a new robot arm and task, fine-tuning only about 1.4% of its parameters with LoRA achieves results comparable to full-parameter fine-tuning.

Also called
OpenVLA-7B, OpenVLA: An Open-Source Vision-Language-Action Model
Related
Vision-Language-Action Model · RT-2 · Open X-Embodiment · Prismatic VLMs · OpenVLA-OFT · LoRA
Sources
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)
OpenVLA 项目主页 (Chinese)
As of
2024-09

See it in the full glossary →