Embodied AI Glossary中文

Qwen-VL

通义千问 Qwen-VLCommon

Alibaba's Qwen team's open-source vision-language model series, frequently used as a backbone for VLA models.

Qwen-VL is the vision-language model series from Alibaba's Qwen team, able to look at images and video and answer in text. The first version was released in August 2023. Qwen2.5-VL came out in early 2025, in sizes from 3B to 72B, supports dynamic-resolution input, and can localize objects with boxes or points. Qwen3-VL began rolling out from September 2025, with dense versions from 2B to 32B and two mixture-of-experts versions, 30B-A3B and 235B-A22B, with native support for a 256K context window. Open weights, small model sizes, and strong grounding ability have made it a popular choice as the vision-language backbone for VLA models. Starting with Qwen3.5 in February 2026, Alibaba's mainline models natively support image input themselves.

ExampleXiaomi's Xiaomi-Robotics-0 uses Qwen3-VL-4B-Instruct as its vision-language backbone, followed by a diffusion Transformer that generates actions; Shanghai AI Lab's InternVLA-M1 uses Qwen2.5-VL-3B as its System 2.

Also called
Qwen2-VL, Qwen2.5-VL, Qwen3-VL
Related
Vision-Language Model · Multimodal Large Language Model · PaliGemma · InternVL · Vision-Language-Action Model · Dynamic / Native Resolution
Sources
QwenLM/Qwen3-VL GitHub (News)
Qwen2.5-VL Technical Report (arXiv 2502.13923)
Xiaomi-Robotics-0 (arXiv 2602.12684)
As of
2026-09

See it in the full glossary →