Embodied AI Glossary中文

Value Function

价值函数Common

A function estimating how much total reward will follow from a given state if the agent keeps following a given policy.

The value function is a core concept in reinforcement learning. The state-value function V(s) is the expected return — the discounted sum of future rewards — from state s onward while following policy π the whole time; the Q-function Q(s,a) is the expected return from taking action a first and then following the policy; A = Q − V is called the advantage function, measuring how much better a given action is than average. With it, an agent can judge which states and actions are worth pursuing. The actor-critic architecture combines a policy with a value function: the actor is the policy network, responsible for producing actions, and the critic is the value network, scoring actions and supplying the advantage signal used to update the actor — PPO, SAC, and TD3 all belong to this family. In embodied AI, value functions are also used to filter data and to run reinforcement learning on VLAs, as in Physical Intelligence's RECAP method for π*0.6.

Exampleπ*0.6 trains a value function that predicts “how many steps remain until the task succeeds” (a failed trajectory gets a very low value), and uses it to compute each action's advantage, telling the policy which actions are better.

Also called
State-Value Function, V-Function, Critic
Related
Q-Function · Advantage Function · Bellman Equation · Temporal-Difference Learning · Return · Asymmetric Actor-Critic
Sources
OpenAI Spinning Up: Key Concepts in RL(Value Functions)
Hugging Face Deep RL Course: Advantage Actor-Critic (A2C)
π*0.6: a VLA That Learns From Experience (arXiv 2511.14759)

See it in the full glossary →