REINFORCE
REINFORCE 算法AdvancedThe earliest policy-gradient algorithm: raise the probability of actions that led to high return across a full trajectory.
REINFORCE was proposed by Ronald Williams in 1992, and is the earliest policy-gradient algorithm. It runs the current policy through several full episodes and, for every action taken, weights the gradient of that action's log-probability by the cumulative return that followed it, so actions that led to higher return get pushed to be more likely. Because the return comes directly from sampling a whole trajectory rather than from a learned value estimate, it is also called the Monte Carlo policy gradient; the estimate is unbiased but high-variance and sample-inefficient. Subtracting a baseline, such as the average return or a value function, noticeably reduces variance, and this idea later grew into actor-critic methods. PPO also builds on policy gradients; in large-language-model post-training, RLOO and GRPO use the average return within a group as a baseline, carrying REINFORCE's basic idea forward.
ExampleTraining an arm to push a block: an episode that reaches the goal scores 1, otherwise 0. REINFORCE raises the probability of every action in a successful episode, while a failed episode gets zero gradient; with a baseline added, actions in below-average episodes get pushed down instead.
- Also called
- Monte Carlo Policy Gradient
- Related
- Policy Gradient · Monte Carlo Methods · Return · Advantage Function · Proximal Policy Optimization · Group Relative Policy Optimization
- Sources
- Wikipedia: Policy gradient method
Ahmadian et al. 2024: Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs