Embodied AI Glossary中文

VLA-0

Advanced

NVIDIA's minimalist VLA that makes no architecture changes at all, having the VLM write out actions directly as plain digit text.

VLA-0 is work by Ankit Goyal and colleagues at NVIDIA, released in October 2025. Typical VLAs (vision-language-action models) either add special action tokens to the vision-language model's vocabulary or attach a separate action head; VLA-0 changes nothing at all, having Qwen2.5-VL-3B output actions in plain text: each continuous action dimension is normalized into an integer between 0 and 1000, written out as digits, and decoded back afterward. The paper points out that making this work well needs a couple of tricks: during training, characters in the target digit string are randomly masked to force the model to actually look at the image rather than just continuing the text; at inference, predictions from adjacent timesteps are ensembled and averaged. The result is 94.7% average success on LIBERO, beating OpenVLA-OFT and SmolVLA when trained only on that benchmark's own data, and also beating models pretrained on large-scale robot data such as π0 and GR00T N1; on a real SO-100 robot it beats SmolVLA by 12.5 points. It is a reminder to get the simple baseline solid before reaching for something fancier.

ExampleGiven a camera image and an instruction, the model answers like a question, outputting a string of integers between 0 and 1000 as text, with each number corresponding to one action dimension; these are un-normalized and handed to the robot arm for execution.

Also called
VLA-0 (Action as Text), VLA-0: Building State-of-the-Art VLAs with Zero Modification
Related
Vision-Language-Action Model · Action Tokenizer · Action Binning · Qwen-VL · LIBERO Benchmark · SmolVLA
Sources
VLA-0: Building State-of-the-Art VLAs with Zero Modification (arXiv 2510.13054)
VLA-0 项目主页 (Chinese)
As of
2025-10

See it in the full glossary →