Embodied AI Glossary中文

Deterministic vs. Stochastic Policy

确定性策略 / 随机策略Advanced

A deterministic policy always gives the same action for the same state; a stochastic policy gives a probability distribution to sample from.

This is a classification of policies (mappings from observation to action) by the form of their output. A deterministic policy is written a = μ(s), always outputting the same action for the same state; a stochastic policy is written a ~ π(·|s), outputting a probability distribution over actions, which may sample differently each time. Discrete actions commonly use a categorical distribution, giving a probability per action via softmax like a classifier; continuous actions commonly use a diagonal Gaussian, with the network outputting a mean and a log standard deviation. In reinforcement learning, stochastic policies carry exploration built in, and PPO and SAC both use them; DDPG and TD3 use deterministic policies and add noise separately during training for exploration, which makes the policy-gradient estimate more efficient. In imitation learning, behavior cloning trained with mean squared error regression is essentially deterministic, and when the same scene in the demonstrations has multiple valid solutions, it gets averaged into one wrong action; stochastic policies like diffusion policies or Gaussian mixture models can represent this action multimodality instead.

ExampleA robot arm reaching around an obstacle to grab a cup, with half the demonstrations going left and half going right: a deterministic policy trained with mean squared error regression outputs the average of the two and drives straight into the obstacle, while a stochastic policy like a diffusion policy lands on either the left or the right route on any given sample.

Also called
Deterministic Policy, Stochastic Policy
Related
Policy · Gaussian Policy · Action Multimodality · Deep Deterministic Policy Gradient · Diffusion Policy · Exploration vs. Exploitation
Sources
OpenAI Spinning Up: Key Concepts in RL (Policies)
Deterministic Policy Gradient Algorithms (ICML 2014)
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)

See it in the full glossary →