Action Tokenizer
动作分词器CommonA module that encodes continuous actions into discrete tokens and decodes them back into actions at inference time.
An action tokenizer converts a continuous action, usually a whole action chunk, into discrete tokens, and decodes them back into an executable action at inference time, letting a VLA output actions the same way it predicts the next token. The simplest version is per-dimension binning, as in RT-2 and OpenVLA's 256 bins per dimension, but with high-frequency data, actions at neighboring timesteps are nearly identical, so binning produces a lot of redundant tokens the model struggles to learn from. FAST, proposed by Physical Intelligence and others in January 2025, instead applies a discrete cosine transform, DCT, converting a signal into the frequency domain, to each action dimension, quantizes it, and compresses it with byte-pair encoding (BPE), cutting π0's training time to about a fifth. Another line of work learns a discrete action codebook using vector quantization (VQ).
ExampleOne second of 50Hz action data for folding a T-shirt needs 700 tokens with per-dimension binning, but only 53 tokens after FAST compression.
- Also called
- Action Tokenization, Action Token
- Related
- Action Binning · Discrete Cosine Transform · Byte-Pair Encoding · Vector Quantization · π0-FAST · Action Representation
- Sources
- FAST: Efficient Action Tokenization for Vision-Language-Action Models (arXiv:2501.09747)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246) - As of
- 2025-01