Embodied AI Glossary中文

Adversarial Motion Priors

对抗运动先验AMPAdvanced

Using a discriminator that judges how much motion resembles motion-capture data, and turning that resemblance into a reward for natural movement.

AMP was introduced by Xue Bin Peng, Pieter Abbeel, Sergey Levine, Angjoo Kanazawa, and colleagues in 2021, originally for simulated character animation. It borrows from generative adversarial imitation learning: a discriminator is trained on pairs of consecutive-frame states, judging whether a snippet of motion came from a reference motion-capture dataset or from the policy; the policy treats how well it fools the discriminator as a style reward, added to the task reward (such as moving forward at a commanded speed) and optimized together with reinforcement learning. The benefit is not having to hand-write an imitation objective or pick specific motion clips to track — an unstructured pile of mocap clips is enough to make the motion look natural. In 2022, Escontrela and colleagues applied it to the Unitree A1 quadruped, learning a natural, energy-efficient gait from only about 4.5 seconds of German shepherd motion-capture data, and it's often used as a substitute for laborious reward engineering.

ExampleEscontrela and colleagues set the style-reward weight to 0.65 and the task-reward weight to 0.35 on the A1; the robot dog learned a gait resembling a real dog's, naturally switching gaits with speed, and transferred directly to the real robot.

Also called
AMP, AMP Style Reward
Related
Generative Adversarial Imitation Learning · Reward Engineering · DeepMimic · ASE · Motion Tracking · Sim-to-Real Transfer
Sources
AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control (arXiv 2104.02180)
Adversarial Motion Priors Make Good Substitutes for Complex Reward Functions (arXiv 2203.15103)

See it in the full glossary →