AMO
AdvancedUCSD's 2025 humanoid whole-body control method, using trajectory optimization to help an RL policy bend and reach far.
AMO was proposed by Xiaolong Wang's group at UC San Diego, published at RSS 2025. Picking something off the floor or reaching a high shelf requires a humanoid to coordinate bending at the waist, twisting the torso, and flexing the legs — but reinforcement-learning policies trained by imitating human motion-capture data rarely see this kind of extreme torso posture, and become unstable on out-of-distribution commands. AMO first uses trajectory optimization (solving for joint trajectories under dynamics constraints) to batch-generate data on “how the legs should move for a given torso orientation and height,” training a small MLP module that supplies a reference pose to the lower-body policy in real time; that lower-body policy is trained in Isaac Gym with teacher-student distillation. The system is deployed on a 29-degree-of-freedom Unitree G1, can be teleoperated with VR, and teleoperated data can also train a Transformer policy to act autonomously.
ExampleIn the paper's demos, a G1 moves cans between surfaces at different heights, taking a bottle from a tall shelf on the left and setting it on a low table on the right, and can also straighten both legs to place a bottle on a high shelf; during teleoperation, 3 poses from the VR device are converted into control commands.
- Also called
- Adaptive Motion Optimization, AMO: Adaptive Motion Optimization for Hyper-Dexterous Humanoid Whole-Body Control
- Related
- Whole-Body Control · Learning-Based Whole-Body Control · Trajectory Optimization · Teacher-Student Distillation · Unitree G1 · Whole-Body Teleoperation
- Sources
- AMO: Adaptive Motion Optimization for Hyper-Dexterous Humanoid Whole-Body Control (arXiv 2505.03738)
AMO 项目页(RSS 2025) (Chinese) - As of
- 2025-05