Embodied AI Glossary中文

Gaussian Policy

高斯策略Advanced

A stochastic policy where the network outputs an action's mean and standard deviation, then samples from that normal distribution.

A Gaussian policy is the most common form of stochastic policy for continuous-action reinforcement learning. A neural network outputs a mean for each action dimension given the observation; the standard deviation is either also output by the network or kept as a separate learnable parameter independent of the observation, usually stored in log form to keep it positive. The action is sampled as 'mean + standard deviation × standard normal noise,' with each dimension usually assumed independent, which is why it's also called a diagonal Gaussian policy. It has a closed-form log-probability, which makes computing the policy gradient easy, and the size of the standard deviation directly controls how much exploration happens. Mainstream algorithms like PPO and SAC both use it, and reinforcement learning for legged locomotion control is almost always built on this kind of policy. At deployment, the mean alone is usually taken as the action. Because it has only one peak, it can't represent several sharply different valid actions, which is why imitation learning often replaces it with a GMM or Diffusion Policy.

ExampleTraining quadruped walking in Isaac Lab with rsl_rl's PPO: the policy outputs the means of 12 target joint angles, paired with a set of learnable standard deviations for exploration during sampling; on the real robot, only the mean is used.

Also called
Diagonal Gaussian Policy
Related
Policy · Deterministic vs. Stochastic Policy · Proximal Policy Optimization · Soft Actor-Critic · Entropy Regularization · Continuous Action Regression
Sources
OpenAI Spinning Up: Key Concepts in RL (Diagonal Gaussian Policies)
leggedrobotics/rsl_rl (GitHub)
Soft Actor-Critic (Haarnoja et al., ICML 2018)

See it in the full glossary →