Embodied AI Glossary中文

History Encoder

历史编码器Advanced

A module that compresses a recent window of observations and actions into a vector, letting the policy infer current conditions.

A history encoder is the part of a policy network dedicated to processing information from the past several steps, commonly implemented with a 1D convolution, an RNN, or a Transformer. The most typical use is in reinforcement learning for legged robots: real-world parameters like ground friction, payload, and motor condition can't be measured directly, but they leave traces in the recent sequence of states and actions. RMA (Kumar et al., RSS 2021) trains in two stages: first, a base policy is trained in simulation using an 8-dimensional extrinsics vector compressed from privileged information (ground-truth parameters only available in simulation); then, an 'adaptation module' is trained to regress that same vector just from the last 50 steps (0.5 seconds) of state-and-action history. At deployment, the adaptation module runs at 10 Hz while the base policy runs at 100 Hz. This 'teacher first, then student' approach has since been widely reused across legged and humanoid locomotion work.

ExampleSuddenly strapping extra weight onto a Unitree A1's back, RMA's adaptation module estimates the new extrinsics from the joint response over the last half second, and the base policy adjusts its gait accordingly in under a second, with no retraining needed.

Also called
Adaptation Module, History Observation Encoder
Related
Rapid Motor Adaptation · Privileged Information · Teacher-Student Distillation · Proprioception · System Identification · Domain Randomization
Sources
RMA: Rapid Motor Adaptation for Legged Robots (Kumar et al., RSS 2021)
RMA paper full text (ar5iv)

See it in the full glossary →