DreamWaQ
AdvancedA 2023 KAIST reinforcement-learning method for blind quadruped locomotion that “imagines” the terrain underfoot using only proprioception.
DreamWaQ is a locomotion control method for quadruped robots from Hyun Myung's group at KAIST (Korea Advanced Institute of Science and Technology), published at ICRA 2023. Many approaches to walking over complex terrain rely on a camera or lidar, but these sensors can fail in bad weather or poor lighting. DreamWaQ instead walks blind, using only proprioception (the robot's own internal signals, such as joint angles, angular velocities, and body orientation): it trains a context-aided estimator network (CENet) that estimates the body's velocity from the last few steps of observation, and uses a variational autoencoder to infer a latent variable representing the terrain — effectively an “implicit imagination” of what's underfoot; the policy then combines these estimates to output joint actions. Training is done in Isaac Gym with 4,096 domain-randomized parallel environments, using an asymmetric actor-critic — where the critic can see privileged information available only in simulation, while the policy sees only what a real robot could observe — before being transferred zero-shot to the real robot.
ExampleOn the Unitree A1 quadruped, the DreamWaQ policy, without any external perception, was able to cross complex terrain such as steps during a single long-distance continuous outdoor walk.
- Also called
- Learning Robust Quadrupedal Locomotion With Implicit Terrain Imagination via Deep Reinforcement Learning
- Related
- Blind Locomotion · Proprioception · Asymmetric Actor-Critic · Privileged Information · Variational Autoencoder · Unitree A1
- Sources
- DreamWaQ (arXiv 2301.10602)
- As of
- 2023-05