Embodied AI Glossary中文

Action Tokenizer

动作分词器Common

A module that encodes continuous actions into discrete tokens and decodes them back into actions at inference time.

An action tokenizer converts a continuous action, usually a whole action chunk, into discrete tokens, and decodes them back into an executable action at inference time, letting a VLA output actions the same way it predicts the next token. The simplest version is per-dimension binning, as in RT-2 and OpenVLA's 256 bins per dimension, but with high-frequency data, actions at neighboring timesteps are nearly identical, so binning produces a lot of redundant tokens the model struggles to learn from. FAST, proposed by Physical Intelligence and others in January 2025, instead applies a discrete cosine transform, DCT, converting a signal into the frequency domain, to each action dimension, quantizes it, and compresses it with byte-pair encoding (BPE), cutting π0's training time to about a fifth. Another line of work learns a discrete action codebook using vector quantization (VQ).

ExampleOne second of 50Hz action data for folding a T-shirt needs 700 tokens with per-dimension binning, but only 53 tokens after FAST compression.

Also called
Action Tokenization, Action Token
Related
Action Binning · Discrete Cosine Transform · Byte-Pair Encoding · Vector Quantization · π0-FAST · Action Representation
Sources
FAST: Efficient Action Tokenization for Vision-Language-Action Models (arXiv:2501.09747)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)
As of
2025-01

See it in the full glossary →