Embodied AI Glossary中文

Goal-Conditioned Behavior Cloning

目标条件模仿学习GCBCAdvanced

Behavior cloning that feeds the policy both the current observation and the goal it's meant to reach.

This builds on ordinary behavior cloning (copying demonstrated actions via supervised learning) by giving the policy one more input: a goal, usually a goal image or goal state, though it can also be language. Training data is often produced through hindsight relabeling: cut a segment out of a longer demonstration and treat its last frame as the goal, so the actions leading up to it automatically become a demonstration of reaching that goal — which means even task-unlabeled “play” data can be used. Google's Lynch and colleagues used it as the baseline Play-GCBC in their 2019 Learning from Play paper; GCSL, from the same year, has the agent collect its own data and repeatedly relabel it before running supervised learning. It's structurally simple and the starting point for many goal-conditioned and language-conditioned policies, but its weak point is that when a single goal has multiple valid routes to it (action multimodality), it tends to learn the average of them instead of any one of them.

ExampleRandomly cutting a short segment out of teleoperated play data and using its last frame as the goal, a policy is trained to output an action given the current image and the goal image.

Also called
GCBC, Goal-Conditioned Imitation Learning
Related
Behavior Cloning · Goal-conditioned Policy · Hindsight Relabeling · Play Data · Goal-Conditioned Reinforcement Learning · Action Multimodality
Sources
Learning Latent Plans from Play(项目页,含 Play-GCBC 基线) (Chinese)
Learning to Reach Goals via Iterated Supervised Learning (GCSL, arXiv 1912.06088)

See it in the full glossary →