Imitation from Observation
从观测中模仿学习IfOAdvancedLearning to imitate from only the demonstrator's states or video, with no recorded action labels at all.
Ordinary imitation learning needs observation-action pairs; imitation from observation gets only a sequence of the demonstrator's states — video of a person doing something, say — with no idea what control command produced each step. This opens the door to using internet video and other huge collections of human video, but it also has to deal with differing viewpoints and body structures. Notable examples: UC Berkeley's Liu and colleagues (2017) used context translation to convert human video into a robot's viewpoint before running reinforcement learning; UT Austin's Torabi and colleagues (2018) proposed BCO, which first lets the agent explore on its own to learn an inverse dynamics model (inferring the action from a pair of consecutive frames), then uses it to fill in action labels for expert video so behavior cloning can be applied; adversarial approaches exist too. Today's embodied-AI work that pretrains on human video, or uses latent actions or pseudo-action labels, is solving exactly this same problem.
ExampleIn Liu and colleagues' 2017 paper, a robot watches only video of a person sweeping, scooping almonds, and pushing objects — no joint recordings at all — and learns to perform the same actions with tools.
- Also called
- IfO, Imitation Learning from Observation, Learning from Video
- Related
- Action-free Video · Inverse Dynamics Model · Human Video Data · Latent Action · Pseudo Action Labels · LAPA
- Sources
- Recent Advances in Imitation Learning from Observation (IJCAI 2019 survey)
Behavioral Cloning from Observation (IJCAI 2018)
Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation (ICRA 2018)