DexVLA
AdvancedA VLA from Midea and others that attaches a roughly billion-parameter diffusion action expert to a vision-language model, adapting to many robots.
DexVLA was released in February 2025 by researchers at Midea Group, East China Normal University, and Shanghai University, published at CoRL 2025. It pairs the vision-language model Qwen2-VL (2 billion parameters) with a pluggable diffusion action expert: the VLM handles looking at images and understanding instructions, while the action expert, built on the ScaleDP architecture and scaled up to about 1 billion parameters, uses multiple output heads to adapt to different robots and is dedicated to generating continuous actions. Training follows an “embodiment curriculum” with three stages: first, the action expert is pretrained on its own using about 100 hours of cross-embodiment data; then it's attached to the VLM and aligned to a specific embodiment; finally, it's post-trained on data annotated with substeps for a specific task. It aims to fix earlier VLAs' weak action representation, expensive training, and difficulty switching robots — the same “VLM plus action expert” approach as π0.
ExampleThe paper tests on a single-arm Franka (gripper or dexterous hand), a bimanual UR5e, and a bimanual AgileX robot; with no task-specific adaptation it folds a shirt (score 0.92), and on the full laundry-folding task it scores 0.4, versus 0.2 for π0 under the same conditions.
- Also called
- DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
- Related
- Vision-Language-Action Model · Action Expert · Diffusion Action Head · Cross-Embodiment · π0 · Qwen-VL
- Sources
- DexVLA (arXiv:2502.05855)
DexVLA 论文 HTML 全文 v3 (Chinese) - As of
- 2025-08