Embodied AI Glossary中文

Discrete Cosine Transform

离散余弦变换DCTAdvanced

A transform that breaks a signal into a weighted sum of cosine waves at different frequencies, widely used for compression.

The discrete cosine transform was originally conceived by Nasir Ahmed, who formally proposed it in a 1974 paper with Natarajan and Rao. It represents a discrete signal as a weighted sum of cosine waves ranging from low to high frequency; those weights are the DCT coefficients. For smooth signals, most of the energy concentrates in a few low-frequency coefficients, so dropping or coarsely quantizing the small high-frequency ones compresses the signal a lot with little loss — this is what JPEG images and MPEG/H.26x video rely on. Embodied AI uses it to represent actions: under high-frequency control, adjacent time steps' actions are very similar, so naively discretizing each step produces a lot of redundant tokens. Physical Intelligence's 2025 FAST tokenizer applies DCT to an action chunk first, then quantizes and compresses it with byte-pair encoding, letting an autoregressive VLA learn high-frequency, dexterous tasks well.

ExampleFAST applies DCT dimension by dimension to an action chunk, quantizes it so most high-frequency coefficients become zero while low-frequency ones are kept, then compresses the result into a handful of tokens with byte-pair encoding; the paper reports that, paired with π0 trained on 10,000 hours of data, this matches the diffusion version's performance while cutting training time up to 5x.

Also called
DCT
Related
Action Tokenizer · π0-FAST · Byte-Pair Encoding · Action Binning · Action Representation · Action Chunking
Sources
Discrete cosine transform - Wikipedia
FAST: Efficient Action Tokenization for Vision-Language-Action Models (arXiv:2501.09747)

See it in the full glossary →