Embodied AI Glossary中文

SmolVLA

Common

Hugging Face's open-source, roughly 450-million-parameter small VLA that trains and runs on ordinary consumer hardware, even a laptop.

SmolVLA is an open-source vision-language-action model, about 450 million parameters, released in June 2025 by Hugging Face's LeRobot team. It's built on the small vision-language model SmolVLM2, followed by an action expert of about 100 million parameters trained with flow matching; to save compute, it keeps only 64 visual tokens per frame and skips some network layers. Pretraining uses only 487 community-uploaded LeRobot datasets, about 10 million frames — an order of magnitude less data than typical VLAs. It also supports asynchronous inference: the robot starts computing the next action chunk while still executing the current one, which the team says speeds up task completion by about 30%. On LIBERO, Meta-World, and real SO-100/SO-101 robots, it performs close to or better than larger models, and it can run on consumer GPUs or even a MacBook.

ExampleCollect a batch of demonstrations with LeRobot of an SO-101 arm picking up blocks, fine-tune SmolVLA on a single consumer GPU, then deploy it back onto the same arm.

Also called
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
Related
LeRobot · Vision-Language-Action Model · Action Expert · Flow Matching · Asynchronous Inference · SO-100 / SO-101 Arm
Sources
SmolVLA: Efficient Vision-Language-Action Model trained on Lerobot Community Data (Hugging Face Blog)
As of
2025-06

See it in the full glossary →