Embodied AI Glossary中文

Behavior Cloning

行为克隆BCEssential

Treating expert demonstrations as labeled data and using supervised learning to directly map observations to actions.

Behavior cloning is the most basic form of imitation learning. You collect a large set of observation-action pairs — for example, the camera images and joint commands recorded while a human teleoperates a robot — treat the observations as inputs and the expert's actions as labels, and train a policy network with ordinary supervised learning to fit them. An early example is Pomerleau's 1988 ALVINN, which used a neural network to output a steering direction directly from images. Behavior cloning is simple and stable, and needs no reward function to be designed, so the core training of ACT, Diffusion Policy, and most VLA (vision-language-action) models is essentially behavior cloning. Its main weakness is compounding error: once the policy drifts even slightly off the demonstrated trajectory, it enters states it never saw during training, and small mistakes snowball from there. DAgger, correction data, and reinforcement-learning fine-tuning are all ways of addressing this weakness.

ExampleACT used behavior cloning on just about 10 minutes — 50 human teleoperated demonstrations — to get a low-cost bimanual robot to succeed 80–90% of the time at fine-grained tasks like opening the lid of a translucent condiment cup or putting batteries into a remote control.

Also called
Behavioral Cloning, Behavioural Cloning, BC
Related
Imitation Learning · Compounding Error · DAgger · Demonstration Data · Action Chunking with Transformers · Diffusion Policy
Sources
Wikipedia: Imitation learning
Pomerleau 1988: ALVINN: An Autonomous Land Vehicle in a Neural Network (NeurIPS)
Zhao et al. 2023: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT)

See it in the full glossary →