Embodied AI Glossary中文

Policy

策略Essential

The rule that decides what action to take next given the current observation — usually a neural network.

Policy is a core concept in reinforcement learning and robot learning: the mapping from the current state or observation to an action, usually written π. A deterministic policy always gives the same action for the same input; a stochastic policy outputs a distribution over actions and samples from it. In embodied AI, a policy is usually a neural network: it takes in camera images, joint angles, and other proprioceptive state, sometimes with a language instruction added, and outputs an end-effector pose, target joint angles, or a walking velocity command. Policies are mainly trained through imitation learning (copying human demonstrations) and reinforcement learning (trial and error guided by reward); VLA models are, at their core, a large-scale form of policy too. Don't confuse this with “model” as used elsewhere in reinforcement learning — there, “model” alone usually means a dynamics model that predicts how the environment changes next (as in “model-based reinforcement learning”), while the policy is what chooses the action.

ExampleDiffusion Policy, on the Push-T task, reads in a camera image and the end effector's current position, and at each step outputs a short upcoming segment of end-effector targets that push a T-shaped block toward the goal position.

Also called
Policy Network
Related
Visuomotor Policy · Imitation Learning · Reinforcement Learning · Observation · Action Space · Vision-Language-Action Model
Sources
OpenAI Spinning Up: Key Concepts in RL
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

See it in the full glossary →