Qwen-VL
通义千问 Qwen-VLCommonAlibaba's Qwen team's open-source vision-language model series, frequently used as a backbone for VLA models.
Qwen-VL is the vision-language model series from Alibaba's Qwen team, able to look at images and video and answer in text. The first version was released in August 2023. Qwen2.5-VL came out in early 2025, in sizes from 3B to 72B, supports dynamic-resolution input, and can localize objects with boxes or points. Qwen3-VL began rolling out from September 2025, with dense versions from 2B to 32B and two mixture-of-experts versions, 30B-A3B and 235B-A22B, with native support for a 256K context window. Open weights, small model sizes, and strong grounding ability have made it a popular choice as the vision-language backbone for VLA models. Starting with Qwen3.5 in February 2026, Alibaba's mainline models natively support image input themselves.
ExampleXiaomi's Xiaomi-Robotics-0 uses Qwen3-VL-4B-Instruct as its vision-language backbone, followed by a diffusion Transformer that generates actions; Shanghai AI Lab's InternVLA-M1 uses Qwen2.5-VL-3B as its System 2.
- Also called
- Qwen2-VL, Qwen2.5-VL, Qwen3-VL
- Related
- Vision-Language Model · Multimodal Large Language Model · PaliGemma · InternVL · Vision-Language-Action Model · Dynamic / Native Resolution
- Sources
- QwenLM/Qwen3-VL GitHub (News)
Qwen2.5-VL Technical Report (arXiv 2502.13923)
Xiaomi-Robotics-0 (arXiv 2602.12684) - As of
- 2026-09