Decoding Strategies
解码策略(贪心 / 温度采样 / Top-k / Top-p)AdvancedThe rule for picking an actual token once the model has given a probability distribution over what comes next.
An autoregressive model only outputs a probability distribution at each step; a decoding strategy decides how to pick from it. Greedy decoding always takes the highest-probability token, giving deterministic but repetition-prone output; beam search keeps several candidate sequences at once; sampling draws randomly according to the probabilities. Temperature adjusts how sharp the distribution is — lower temperature pushes it toward greedy, higher makes it more random. Top-k only samples from the k highest-probability tokens, used by Fan and colleagues for story generation in 2018; top-p (nucleus sampling) samples from the smallest set of tokens whose cumulative probability just exceeds p, proposed by Holtzman and colleagues in 2019 to reduce the dull repetition that maximization-based decoding tends to produce. The same model can produce very different output quality depending on which strategy is used. VLAs that output action tokens usually turn off random sampling at deployment, so actions stay consistent.
ExampleOpenVLA's official example calls predict_action with do_sample=False, meaning action tokens are decoded greedily, so the same image and instruction always produce the same action.
- Also called
- Nucleus Sampling, Top-k Sampling, Top-p Sampling, Beam Search
- Related
- Autoregressive Decoding · Large Language Model · Action Tokenizer · Best-of-N Sampling · Inference-Time Compute · Softmax
- Sources
- How to generate text: using different decoding methods for language generation with Transformers (Hugging Face)
The Curious Case of Neural Text Degeneration (arXiv:1904.09751)
openvla/openvla-7b model card