Continuous Action Regression
连续动作回归CommonHaving a network output continuous action values directly, trained with L1 or mean-squared error against demonstrations.
Continuous action regression is the most direct way for a robot policy to output actions: the network ends in a regression head, usually a few fully connected layers, that directly outputs real numbers such as joint angles or end-effector displacement, trained with mean squared error or L1 loss, the squared or absolute difference between the prediction and the demonstrated action, to match human demonstrations. It is simple to implement, and inference needs just one forward pass, so it's fast. Its weakness shows up when a demonstration set contains several valid ways to do the same thing in the same scene, called action multimodality: regression tends to learn the average of those ways, which can end up being none of them; the Octo paper observed that an MSE regression head produces hesitant, indecisive robot motion. Common alternatives are discretizing actions into tokens, or generating them with diffusion or flow matching. Still, 2025's OpenVLA-OFT paired L1 regression with action chunking and parallel decoding to raise LIBERO's average success rate across four task suites from 76.5% to 97.1%, showing regression is good enough for plenty of tasks.
ExampleACT (Action Chunking with Transformers) trains by directly regressing a future sequence of joint angles with an L1 loss; OpenVLA-OFT replaces OpenVLA's discrete-token output with an L1 regression head, computing a whole action segment in one forward pass and increasing action-generation throughput by about 26x.
- Also called
- Regression Head, L1 Regression, MSE Regression
- Related
- Action Head · Action Multimodality · L1 Loss · Mean Squared Error · Diffusion Action Head · Action Binning
- Sources
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT, arXiv 2502.19645)
Octo: An Open-Source Generalist Robot Policy (arXiv 2405.12213) - As of
- 2025-02