One-shot Imitation Learning
单样本模仿学习AdvancedThe robot watches a new task demonstrated just once, then completes it starting from a different arrangement.
One-shot imitation learning requires the robot, when faced with a new task, to succeed in a new situation — different object placement, different initial state — after seeing just a single demonstration, whether a teleoperated trajectory or a video. Rather than training from scratch on that one demonstration, the approach first trains a policy conditioned on a demonstration across many tasks: feed it a demonstration plus the current observation, and it outputs an action. OpenAI's Yan Duan and colleagues proposed and named this setting in 2017 on a block-stacking task; that same year Chelsea Finn and colleagues used MAML for meta-imitation learning, extending it to raw image input and verifying it on a real robot. Later work tried using human video directly as the demonstration, which also has to cross the embodiment gap between a human and a robot. In essence this is meta-learning applied to imitation learning, and it is conceptually close to large models' in-context learning.
ExampleIn Duan and colleagues' experiments, each task is stacking blocks on a table in some particular pattern, say all into one tower, or into several two-block towers; at test time, given one demonstration of a new pattern, the network has to reproduce that pattern with the blocks starting in new positions.
- Also called
- One-shot Imitation
- Related
- Meta-Learning · Imitation Learning · Few-shot · In-Context Learning · Demonstration Data · OKAMI
- Sources
- Duan et al. 2017: One-Shot Imitation Learning
Finn et al. 2017: One-Shot Visual Imitation Learning via Meta-Learning