Embodied AI Glossary中文

Action Chunking

动作分块Essential

A policy predicts a short sequence of upcoming actions at once, instead of outputting just one action per step.

Action chunking means a policy's inference outputs a sequence of the next k steps of action at once, a “chunk,” executes some or all of it, then re-observes and predicts again. The term entered robot learning through Tony Zhao, Chelsea Finn, and colleagues' 2023 ACT paper, borrowed from a neuroscience idea about packaging a sequence of movements into one unit for execution: ACT controls a dual-arm robot at 50Hz, predicting the next 100 steps of joint targets each time. It solves two problems: it cuts the number of decisions to 1/k of the original, easing the compounding error that behavior cloning suffers from small per-step mistakes; and it makes it easier to learn timing habits in human demonstrations, such as pauses, that a single-step policy struggles to model. Mainstream models today — Diffusion Policy, π0 (50 steps at a time), GR00T N1 (16 steps at a time) — all output action chunks, usually stitched together across chunks with temporal ensembling or real-time chunking to avoid sudden jumps in motion.

ExampleACT performed six fine-motor tasks on the low-cost dual-arm platform ALOHA, each using only 10 to 20 minutes (about 50 demonstrations) of data: opening a translucent condiment-cup lid succeeded 84% of the time, inserting a battery into a remote 96%, and the hardest task, threading a velcro strap, only 20%. It outputs the next 100 steps of joint targets each time, about 2 seconds at 50Hz.

Also called
Action Chunk, Chunk, Action Sequence Prediction
Related
Action Horizon · Temporal Ensembling · Real-Time Chunking · Compounding Error · Action Chunking with Transformers · Diffusion Policy
Sources
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)

See it in the full glossary →