Embodied AI Glossary中文

TinyVLA

Advanced

A fast VLA that pairs a small multimodal model with a diffusion policy head and skips robot-data pretraining altogether.

TinyVLA was released by Midea Group's AI Lab together with East China Normal University and others (Junjie Wen, Yichen Zhu, and colleagues) in September 2024, published in IEEE RA-L 2025. Models like OpenVLA at the time were slow at inference and also required pretraining on large-scale robot data first. TinyVLA takes a different approach: it first trains a family of small vision-language models, ranging from 70 million to 1.4 billion parameters, using Pythia as the language model and following the LLaVA training recipe, to serve as the policy backbone; when fine-tuning on robot data, the pretrained part is frozen and only about 5% of the parameters are trained via LoRA, with a diffusion policy decoder attached to output continuous actions directly. On real single-arm Franka and dual-arm UR5 robots, the largest variant, TinyVLA-H, beats OpenVLA's success rate by 25.7 points while using 5.5x fewer parameters and roughly 20x lower inference latency. The same team later built DexVLA.

ExampleOn a dual-arm UR5 task, OpenVLA — which relies on pretraining with single-arm Open X-Embodiment data — struggles, while TinyVLA, which skips robot pretraining and fine-tunes directly, performs better.

Also called
Tiny-VLA, TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
Related
Vision-Language-Action Model · OpenVLA · Diffusion Policy · LoRA · Inference Latency · DexVLA
Sources
TinyVLA (arXiv 2409.12514)
TinyVLA project page
As of
2025-05

See it in the full glossary →