Embodied AI Glossary中文

Parallel Decoding

并行解码Advanced

Producing an entire sequence in one forward pass, instead of generating it token by token like autoregressive decoding.

Parallel decoding is defined in contrast to autoregressive decoding: an autoregressive model only predicts the next token each time, so producing N tokens takes N forward passes; parallel decoding instead has the model predict every position in a single forward pass. It was first studied systematically in machine translation as 'non-autoregressive translation' (Gu et al., 2017), cutting latency by roughly an order of magnitude at the cost of weaker modeling of dependencies between positions and somewhat lower quality. The issue is especially visible in VLAs: OpenVLA discretizes each action dimension into a single token and generates them one at a time, so a 7-dimensional action needs 7 forward passes, and action chunking makes it even slower. OpenVLA-OFT instead feeds in a set of empty action placeholder embeddings and replaces the causal attention mask with bidirectional attention, producing the entire action chunk in a single forward pass. Methods like discrete diffusion also fall under parallel or few-step parallel decoding.

ExampleOpenVLA-OFT combines parallel decoding with action chunking (outputting 8 steps of 7-dimensional action at once) on LIBERO, raising action throughput from OpenVLA's 4.2 Hz to 108.8 Hz, about 26x faster.

Also called
Non-autoregressive Decoding
Related
Autoregressive Decoding · Action Chunking · OpenVLA-OFT · Causal Attention · Discrete Diffusion · Speculative Decoding
Sources
Non-Autoregressive Neural Machine Translation (arXiv:1711.02281)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT, arXiv:2502.19645)

See it in the full glossary →