Generative Adversarial Imitation Learning
生成对抗模仿学习GAILAdvancedUsing a discriminator to tell expert actions from policy actions, forcing the policy to act more and more like the expert.
Introduced by Stanford's Jonathan Ho and Stefano Ermon in 2016. The traditional route was two separate steps: recover the expert's reward function with inverse reinforcement learning (working backward from demonstrations), then train a policy with reinforcement learning on that reward — slow and indirect. GAIL borrows the structure of a generative adversarial network instead: a discriminator learns to tell whether a state-action pair came from an expert demonstration or from the current policy, and the policy treats the discriminator's output as its reward, updating with TRPO (a policy-gradient algorithm that limits how much the policy can change per step) until the discriminator can no longer tell the difference. It requires repeated interaction with the environment, so it's mostly used in simulation, but it's less prone to compounding error (small mistakes snowballing) than behavior cloning. Adversarial motion priors (AMP), commonly used for humanoid robots and character animation, grew out of this same adversarial-imitation idea.
ExampleThe original paper, given only a handful of expert trajectories and no reward function, trains GAIL policies on MuJoCo-simulated tasks like humanoid walking that mostly reach over 70% of expert-level performance.
- Also called
- GAIL
- Related
- Inverse Reinforcement Learning · Imitation Learning · Behavior Cloning · Generative Adversarial Network · Adversarial Motion Priors · Trust Region Policy Optimization
- Sources
- Generative Adversarial Imitation Learning (Ho & Ermon, arXiv 1606.03476)