Parallel Decoding
并行解码AdvancedProducing an entire sequence in one forward pass, instead of generating it token by token like autoregressive decoding.
Parallel decoding is defined in contrast to autoregressive decoding: an autoregressive model only predicts the next token each time, so producing N tokens takes N forward passes; parallel decoding instead has the model predict every position in a single forward pass. It was first studied systematically in machine translation as 'non-autoregressive translation' (Gu et al., 2017), cutting latency by roughly an order of magnitude at the cost of weaker modeling of dependencies between positions and somewhat lower quality. The issue is especially visible in VLAs: OpenVLA discretizes each action dimension into a single token and generates them one at a time, so a 7-dimensional action needs 7 forward passes, and action chunking makes it even slower. OpenVLA-OFT instead feeds in a set of empty action placeholder embeddings and replaces the causal attention mask with bidirectional attention, producing the entire action chunk in a single forward pass. Methods like discrete diffusion also fall under parallel or few-step parallel decoding.
ExampleOpenVLA-OFT combines parallel decoding with action chunking (outputting 8 steps of 7-dimensional action at once) on LIBERO, raising action throughput from OpenVLA's 4.2 Hz to 108.8 Hz, about 26x faster.
- Also called
- Non-autoregressive Decoding
- Related
- Autoregressive Decoding · Action Chunking · OpenVLA-OFT · Causal Attention · Discrete Diffusion · Speculative Decoding
- Sources
- Non-Autoregressive Neural Machine Translation (arXiv:1711.02281)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT, arXiv:2502.19645)