CogACT
AdvancedAn open-source VLA that attaches a diffusion Transformer action module behind a frozen vision-language backbone.
CogACT is a vision-language-action model released in November 2024 by Microsoft Research Asia together with Tsinghua University, the University of Science and Technology of China, and others. At the time, models like OpenVLA had the VLM output discrete action tokens directly, limiting precision and continuity. CogACT separates cognition from action: a VLM such as Prismatic-7B understands the image and instruction and outputs a cognitive feature, and a diffusion Transformer (DiT) action module of up to about 300 million parameters then generates a continuous action sequence conditioned on that feature. At inference, adaptive action ensembling averages only similar action predictions together, avoiding mixing different modes. Trained on about 400,000 trajectories from Open X-Embodiment, it beats OpenVLA by more than 35% in SIMPLER simulation and 55% on real robots, and even outperforms the 55-billion-parameter RT-2-X in simulation.
ExampleResearchers tested CogACT on two different real arms, a Realman and a Franka, and it still completed manipulation tasks with objects and backgrounds it had never seen.
- Also called
- CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
- Related
- Vision-Language-Action Model · Diffusion Action Head · Action Expert · OpenVLA · Prismatic VLMs · Temporal Ensembling
- Sources
- CogACT (arXiv 2411.19650)
CogACT 项目主页 (Chinese) - As of
- 2024-11