Discount Factor
折扣因子γCommonThe coefficient in reinforcement learning that discounts future rewards, controlling how much the agent weighs long-term payoff.
The discount factor is a basic reinforcement-learning hyperparameter, written γ, with a value between 0 and 1. When computing the return (the sum of rewards accumulated from the current moment onward), a reward k steps in the future gets multiplied by γ raised to the k-th power. It serves two purposes: expressing that a reward received sooner is more certain and more valuable than one received later, and turning an infinitely long sum of rewards into a finite, well-behaved value, which makes it solvable with the Bellman equation (the recursive relationship that defines a value function). A γ closer to 1 makes the agent weigh long-term outcomes more heavily, but also makes the value estimate harder to learn and noisier; a smaller γ makes the agent more shortsighted. Robot reinforcement learning commonly uses a γ around 0.99. It appears in the definitions of return, the value function, and the advantage function.
Examplelegged_gym (ETH's open-source reinforcement-learning framework for training legged robots) sets gamma = 0.99 in its PPO config: a reward 100 steps away is weighted by roughly 0.99 to the 100th power, about 0.37.
- Also called
- γ (gamma), Discount Rate
- Related
- Return · Reward Function · Value Function · Bellman Equation · Generalized Advantage Estimation · Markov Decision Process
- Sources
- OpenAI Spinning Up: Key Concepts in RL
legged_gym: legged_robot_config.py (PPO gamma = 0.99)