Embodied AI Glossary中文

Reinforcement Learning

强化学习RLEssential

A machine-learning approach where an agent improves its behavior through repeated trial and error, guided by a reward signal.

Reinforcement learning is a machine-learning paradigm alongside supervised and unsupervised learning. An agent observes a state in an environment, takes an action, and the environment returns a new state along with a reward (a number measuring how good that step was); the goal is to learn a policy that maximizes long-term cumulative reward. It doesn't require labeling the “correct action” at every step — only a reward needs to be defined — but the agent has to balance trying new actions (exploration) against using actions already known to be good (exploitation). The problem is usually formalized as a Markov decision process. In embodied AI, reinforcement learning is mainly used in two places: training legged and humanoid robot locomotion control at massive scale in parallel simulation, then transferring it to the real robot; and fine-tuning a VLA model with reinforcement learning after an imitation-learning base has been trained, to improve success rate and error recovery.

ExampleIn 2018, OpenAI trained a Shadow dexterous hand entirely in simulation using reinforcement learning, randomizing physical parameters like friction coefficients; the resulting policy transferred directly to a real dexterous hand and performed vision-based in-hand object reorientation.

Related
Markov Decision Process · Reward Function · Policy · Proximal Policy Optimization · Sim-to-Real Transfer · Reinforcement Fine-Tuning (RL Fine-Tuning)
Sources
Wikipedia: Reinforcement learning
OpenAI Spinning Up: Key Concepts in RL
OpenAI et al. 2018: Learning Dexterous In-Hand Manipulation

See it in the full glossary →