Embodied AI Glossary中文

Q-Function

Q 函数Common

The expected return of taking a specific action in a state and then following the policy from then on.

The Q-function is written Q(s,a): the expected return — the discounted sum of future rewards — from taking action a in state s and then following policy π forever after. Compared with the value function V(s), it takes an extra action input, so it can directly compare how good different actions are in the same state; once the optimal Q-function is known, picking the action with the highest Q-value at every step is the optimal policy. The Q-function satisfies the Bellman equation: the current Q-value equals the immediate reward plus the discounted Q-value of what follows. Q-learning, DQN, SAC, and TD3 all learn it; the critic in an actor-critic algorithm is often literally a Q-network, and the advantage function A(s,a) = Q(s,a) − V(s) is derived from it too.

ExampleQT-Opt uses a convolutional network to estimate Q-values: it takes the current camera image and a candidate gripper action as input and outputs the probability that this action ultimately leads to a successful grasp; at every step it searches for the action with the highest Q-value using a cross-entropy method.

Also called
Action-Value Function, Q-value, Q(s,a)
Related
Value Function · Advantage Function · Bellman Equation · Q-Learning · Deep Q-Network · Return
Sources
OpenAI Spinning Up: Key Concepts in RL
Hugging Face Deep RL Course: Introducing Q-Learning

See it in the full glossary →