Action Chain-of-Thought
动作思维链ACoTAdvancedHaving a VLA first reason out a coarse trajectory in action space, then generate the fine-grained action from it.
Chain-of-thought originally meant having a large model write out intermediate reasoning steps before giving an answer. Existing reasoning approaches in VLAs mostly predict subtask text or generate a goal image as the intermediate step — neither of which is an action itself. In January 2026, a team from Beihang University and AgiBot proposed ACoT-VLA (accepted to CVPR 2026), arguing for reasoning directly in action space: an Explicit Action Reasoner (EAR) uses flow matching to first generate a coarse reference trajectory; an Implicit Action Reasoner (IAR) uses learnable queries to extract a latent action prior from the VLM's internal features; the two are then combined via cross-attention to guide the action head in producing the final action sequence. The paper reports a 98.5% average success rate on LIBERO. It can be seen as another form of reasoning alongside embodied chain-of-thought and visual chain-of-thought. Separately, some 2026 navigation work also uses 'Action-CoT' to refer to step-by-step action reasoning.
ExampleACoT-VLA reports a 66.7% average success rate across three manipulation tasks on a real AgiBot G1 robot, and an 88.0% average success rate on LIBERO-Plus after supervised fine-tuning.
- Also called
- ACoT, ACoT-VLA
- Related
- Chain-of-Thought · Embodied Chain-of-Thought · Visual Chain-of-Thought · Vision-Language-Action Model · Action Head · Flow Matching
- Sources
- ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models (arXiv 2601.11404)
ACoT-VLA 论文 HTML 版 (Chinese) - As of
- 2026-01