Decoder-only Architecture
仅解码器架构AdvancedA Transformer design that keeps only the decoder, predicting the next token one at a time with causal attention.
Decoder-only architecture removes the encoder from the original Transformer and stacks only decoder blocks. Google's Liu and colleagues used it to handle very long sequences in their 2018 work generating Wikipedia articles, and OpenAI's GPT-1 adopted the same design that same year; since then, the GPT series, Llama, and most other mainstream large language models have been decoder-only. It does next-token prediction using causal attention, where each position can only see the tokens before it, giving a simple, unified training objective that scales easily. Multimodal models feed it images by encoding them as visual tokens and placing them ahead of the text. Many VLAs reuse this language-model backbone directly, treating actions as just another kind of output token.
ExampleOpenVLA is built on Llama 2 7B, a decoder-only language model: image features and the text instruction are concatenated into one sequence as input, and the model outputs discretized action tokens one at a time.
- Also called
- Decoder-only, Decoder-only Transformer
- Related
- Transformer · Encoder-Decoder · Causal Attention · Autoregressive Decoding · Large Language Model · Next-Token Prediction
- Sources
- Generating Wikipedia by Summarizing Long Sequences (arXiv:1801.10198)
Generative pre-trained transformer - Wikipedia