starVLA
AdvancedAn open-source codebase for assembling and training VLA models like building blocks.
starVLA is an open-source codebase for developing vision-language-action (VLA) models, self-described as “LEGO-style”: it splits a vision-language-model backbone and various types of action head into swappable modules, so researchers can combine autoregressive discrete-action, continuous-regression, or flow-matching action heads within the same codebase and compare them under one unified training and evaluation pipeline. According to its project page, it uses the Qwen-VL model family as its main backbone and plugs into common benchmarks such as LIBERO and SimplerEnv. It suits newcomers who want to quickly reproduce and compare different VLA design choices.
ExampleKeep the Qwen-VL backbone fixed and swap only the action head, from discrete tokens to a flow-matching head, then compare the two on LIBERO success rate.
- Related
- Vision-Language-Action Model · Action Head · Qwen-VL · LIBERO Benchmark · Dexbotic (Dexmal VLA toolbox) · openpi (Physical Intelligence)
- Sources
- starVLA/starVLA (GitHub)
- As of
- 2026-09