Temporal Ensembling
时序集成CommonPredicting an action chunk at every step, then averaging the overlapping chunks' predictions for the same moment before executing.
Temporal ensembling is a way of executing action chunks introduced in the 2023 ACT paper (Tony Zhao and colleagues). Naive action chunking only looks at a new observation every k steps, so the action can jump abruptly when it switches to a new chunk. Temporal ensembling instead calls the policy at every control step, so at any given moment there are several overlapping chunks' predictions for it; these are combined with exponential weights w_i = exp(-m·i), where i=0 is the earliest prediction, and the weighted average is what actually gets executed — a smaller m lets new observations take effect faster. It adds no training cost, only extra inference compute, and in ACT's ablations it improved success rate by about 3.3 percentage points. The cost is that it requires an inference call at every step, which is hard to afford when a large model has high latency; the later Real-Time Chunking (RTC) method uses it as a comparison baseline.
ExampleACT predicts the next k steps of action at every step; at time t, it has already received several predictions for t from the calls made at t, t-1, t-2, and so on, and averages them before sending the result to the arm, which is why ALOHA's bimanual motion looks smoother.
- Also called
- Temporal Ensemble, Action Ensemble
- Related
- Action Chunking · Action Chunking with Transformers · Real-Time Chunking · Action Smoothing · Compounding Error · Action Horizon
- Sources
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)
Real-Time Execution of Action Chunking Flow Policies (arXiv 2506.07339)