DexGraspVLA
AdvancedA hierarchical dexterous-hand grasping framework from Peking University and PsiBot: a large model plans, a diffusion policy acts.
DexGraspVLA was released in February 2025 by teams including Peking University's Institute for Artificial Intelligence and the Peking University–PsiBot (灵初智能) joint lab, later accepted as an oral presentation at AAAI 2026. It has two layers. The upper layer uses an off-the-shelf vision-language model (such as Qwen-VL) to understand the instruction and draw a box around the target object in the image. The lower layer is a diffusion controller: SAM and Cutie continuously track the target's mask, a frozen DINOv2 extracts visual features, and a diffusion Transformer outputs arm and dexterous-hand actions. The core idea is to use foundation models first to turn highly variable images and language into a more stable representation, narrowing the gap between training and test scenes, so that imitation learning from only a small number of human demonstrations can still generalize to a huge number of new scenes.
ExampleAfter training on 2,094 cluttered-scene grasping demonstrations collected across 36 household objects, it was tested zero-shot across roughly 1,287 scenes combining 360 new objects, 6 new backgrounds, and 3 new lighting conditions, reaching an overall success rate of 90.8% on a Realman 7-DOF arm paired with PsiBot's 6-DOF G0-R dexterous hand.
- Also called
- DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping
- Related
- Vision-Language-Action Model · Dexterous Manipulation · Hierarchical Architecture · Diffusion Policy · DINOv2 · PsiBot
- Sources
- DexGraspVLA (arXiv:2502.20900)
DexGraspVLA 项目主页 (Chinese) - As of
- 2025-11