Q-Chunking
动作分块强化学习QCAdvancedDoing reinforcement learning where both the policy and the Q-function operate on a whole chunk of actions at once.
Q-chunking is a reinforcement learning method proposed by Sergey Levine's team at Berkeley in 2025, aimed at long-horizon, sparse-reward, offline-to-online tasks, where a policy is pretrained on existing data and then goes online. It brings action chunking, a technique common in imitation learning, into reinforcement learning: the policy outputs a chunk of h consecutive actions at once, and the Q-function, which judges how good an action is, also takes the state plus the whole action chunk as input. This has two benefits: acting in chunks makes exploration more coherent and lets the agent carry over behavioral habits from the offline data, and scoring a whole chunk at once enables unbiased multi-step temporal-difference updates, so value information propagates faster. To keep the policy from drifting too far from the data, it samples several action chunks from a flow-matching policy and picks the one with the highest Q-value, or adds a distillation constraint.
ExampleOn OGBench's block- and puzzle-manipulation tasks and on robomimic tasks, with 1 million steps of offline pretraining followed by 1 million steps of online interaction, Q-chunking with a chunk length of 5 clearly beats offline-to-online baselines such as RLPD, and the harder the task, the bigger the gap.
- Also called
- Reinforcement Learning with Action Chunking, QC
- Related
- Action Chunking · Offline-to-Online Reinforcement Learning · Temporal-Difference Learning · Reinforcement Learning with Prior Data · Flow Matching · Sparse Reward
- Sources
- Li, Zhou, Levine 2025: Reinforcement Learning with Action Chunking
- As of
- 2025-07