Embodied AI Glossary中文

Spirit v1.5

千寻 Spirit v1.5Advanced

An open-source VLA foundation model from Spirit AI, built on the premise that messy, unscripted data makes a better pretraining set.

Spirit v1.5 is a vision-language-action (VLA) model released and open-sourced by Chinese startup Spirit AI (千寻智能) in January 2026, with the inference code (MIT license), base weights, and one fine-tuned checkpoint (Apache 2.0 license) made public, followed by fine-tuning code in April. Its architecture is the common 'VLM plus action head' pattern: Qwen3-VL-4B as the vision-language backbone, followed by a diffusion Transformer (DiT) action head that generates continuous actions. Its central claim, stated in the title of its technical blog post, is that 'clean data is the enemy of a good robot foundation model': rather than giving data collectors a fixed script or staging objects carefully, they are simply given a goal and left to complete a chain of real tasks freely, so the resulting data naturally contains failed retries and task switching. The team states this raised effective per-collector data-collection time by 200%. At release, it ranked first on RoboChallenge's Table30 real-robot leaderboard.

ExampleA data collector sets themselves a goal, such as 'use the robot to mix a drink today,' and the whole session — preparing ingredients, noticing a wrong ratio, adjusting, adding or removing ingredients — is recorded continuously and used for pretraining.

Also called
Spirit-v1.5, Spirit AI Spirit v1.5
Related
Spirit AI · Vision-Language-Action Model · RoboChallenge · Qwen-VL · Diffusion Transformer · Data Diversity
Sources
Spirit-v1.5: Clean Data Is the Enemy of Great Robot Foundation Models (Spirit AI Blog)
Spirit-AI-Team/spirit-v1.5 (GitHub)
Spirit-AI-robotics/Spirit-v1.5 (Hugging Face)
As of
2026-04

See it in the full glossary →