Embodied AI Glossary中文

Inverse Reinforcement Learning

逆强化学习IRLCommon

Working backward from an expert's demonstrated behavior to infer the reward function it's implicitly optimizing.

Inverse reinforcement learning runs reinforcement learning in reverse: ordinary RL starts from a given reward function (the rule that scores behavior) and learns a policy, while IRL instead observes expert demonstrations and infers what reward the expert must be optimizing. Stuart Russell posed this problem in 1998, Andrew Ng and Russell gave the first algorithms in 2000, and Abbeel and Ng applied it to “apprenticeship learning” in 2004: infer the reward first, then use reinforcement learning to find a policy from it. Its value is that many tasks have rewards that are hard to hand-write, and a learned reward often transfers to new environments better than copying actions directly does. The difficulty is that the same behavior can be explained by many different reward functions, so extra assumptions such as maximum entropy are needed to resolve the ambiguity. Later methods like generative adversarial imitation learning (GAIL) and adversarial motion priors (AMP) continue this line of thinking.

ExampleAbbeel, Coates, and Ng used apprenticeship learning to teach an autonomous helicopter aerobatic maneuvers — flips, loops, autorotation landings — by learning from a human pilot's demonstrations, instead of hand-writing a reward function.

Also called
IRL
Related
Reinforcement Learning · Reward Function · Imitation Learning · Generative Adversarial Imitation Learning · Adversarial Motion Priors · Behavior Cloning
Sources
Apprenticeship learning(Wikipedia,含 Inverse reinforcement learning 一节) (Chinese)
A Survey of Inverse Reinforcement Learning: Challenges, Methods and Progress (arXiv 1806.06877)

See it in the full glossary →