Embodied AI Glossary中文

CogACT

Advanced

An open-source VLA that attaches a diffusion Transformer action module behind a frozen vision-language backbone.

CogACT is a vision-language-action model released in November 2024 by Microsoft Research Asia together with Tsinghua University, the University of Science and Technology of China, and others. At the time, models like OpenVLA had the VLM output discrete action tokens directly, limiting precision and continuity. CogACT separates cognition from action: a VLM such as Prismatic-7B understands the image and instruction and outputs a cognitive feature, and a diffusion Transformer (DiT) action module of up to about 300 million parameters then generates a continuous action sequence conditioned on that feature. At inference, adaptive action ensembling averages only similar action predictions together, avoiding mixing different modes. Trained on about 400,000 trajectories from Open X-Embodiment, it beats OpenVLA by more than 35% in SIMPLER simulation and 55% on real robots, and even outperforms the 55-billion-parameter RT-2-X in simulation.

ExampleResearchers tested CogACT on two different real arms, a Realman and a Franka, and it still completed manipulation tasks with objects and backgrounds it had never seen.

Also called
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
Related
Vision-Language-Action Model · Diffusion Action Head · Action Expert · OpenVLA · Prismatic VLMs · Temporal Ensembling
Sources
CogACT (arXiv 2411.19650)
CogACT 项目主页 (Chinese)
As of
2024-11

See it in the full glossary →