π0-FAST
CommonAn autoregressive VLA that tokenizes π0's actions into discrete tokens with the FAST tokenizer and predicts them one by one.
π0-FAST is a model Physical Intelligence released in January 2025 alongside the FAST action tokenizer; it shares π0's backbone network and training data, differing only in how actions are output. π0 generates continuous actions with flow matching, while π0-FAST turns actions into discrete tokens and predicts them one at a time, like a language model. Traditional discretization — binning each dimension at each timestep separately — barely learns anything on high-frequency, dexterous tasks. FAST instead applies a discrete cosine transform (DCT, the same transform used in JPEG compression) to a chunk of actions, quantizes it, and compresses it further with byte-pair encoding (BPE), so a chunk of actions typically needs only 30–60 tokens. This trains up to 5x faster with performance comparable to the flow-matching version, and produced the first generalist policy trained on the DROID dataset able to follow instructions zero-shot in new environments. The cost is that autoregressive decoding is noticeably slower than flow matching. Weights are open-sourced in openpi.
ExampleA bimanual robot controlled at 50Hz with 14 dimensions per step means 700 numbers per second of action. The old per-dimension binning approach would need 700 tokens; after FAST compression, only a few dozen remain, letting the model learn to fold clothes with next-token prediction, then convert tokens back into continuous actions at inference time.
- Also called
- pi0-FAST, pi0_fast, π0 + FAST
- Related
- π0 · Action Tokenizer · Discrete Cosine Transform · Byte-Pair Encoding · Action Binning · Autoregressive Decoding
- Sources
- FAST: Efficient Action Tokenization for Vision-Language-Action Models (arXiv 2501.09747)
FAST: Efficient Robot Action Tokenization (Physical Intelligence)
openpi (GitHub, Physical Intelligence) - As of
- 2025-01