Embodied AI Glossary中文

Implicit Q-Learning

隐式 Q 学习IQLAdvanced

An offline reinforcement-learning algorithm that only ever values actions already in the dataset, never querying an action it hasn't seen.

Introduced by UC Berkeley's Kostrikov, Nair, and Levine in 2021. Offline reinforcement learning trains only on a fixed dataset, and its central difficulty is that the Q-function (an action's estimated value) tends to wildly overestimate actions absent from the data (extrapolation error), and the policy collapses as soon as it chases them. IQL's solution is to never evaluate a new action at all: it first uses expectile regression (a form of regression that leans toward the upper quantiles) to fit a state value V, from the dataset's own actions, that approximates the value of a close-to-the-best action; it then uses that V to do a temporal-difference update of Q; and finally extracts the policy with advantage-weighted regression (weighting behavior cloning by the size of the advantage). It's simple to implement, performs strongly on the D4RL offline benchmark, suits pretraining offline and then fine-tuning online, and is commonly used as a baseline in robot offline-reinforcement-learning work.

ExampleGiven a batch of historical robot-arm data mixing good and bad manipulation attempts, with no further environment interaction, IQL trains a policy that ends up better than the average quality of the data itself.

Also called
IQL
Related
Offline Reinforcement Learning · Advantage-Weighted Regression · Conservative Q-Learning · Extrapolation Error (OOD Actions in Offline RL) · Offline-to-Online Reinforcement Learning · Q-Function
Sources
Offline Reinforcement Learning with Implicit Q-Learning (arXiv 2110.06169)

See it in the full glossary →