Embodied AI Glossary中文

π0

Essential

Physical Intelligence's 2024 VLA that generates continuous robot actions with flow matching; code and weights are open-sourced.

π0 is a general-purpose robot foundation model released on October 31, 2024, by the US embodied-AI company Physical Intelligence (PI). It's built on Google's 3-billion-parameter PaliGemma vision-language model, plus an “action expert” of about 300 million parameters that uses flow matching — a method similar to diffusion models that generates continuous values by gradually denoising — to produce the next 50 steps of action in one shot, controlling at up to 50Hz. Unlike RT-2 or OpenVLA, which discretize actions into tokens, this continuous generation suits high-frequency, dexterous tasks like folding laundry better. It was pretrained on more than 10,000 hours of robot data spanning 7 robot embodiments and 68 tasks, mixed with open datasets like OXE and DROID; after pretraining it can follow language instructions directly, and can also be fine-tuned on a small amount of data to learn new skills. PI later open-sourced the code and weights in the openpi repository, and it can run inference and LoRA fine-tuning on a single RTX 4090.

ExampleIn the paper's demos, a fine-tuned π0 takes clothes out of a dryer, carries them to a table, and folds each one; it can also assemble cardboard boxes and clear a table.

Also called
pi-zero, pi0, π-zero, π0: A Vision-Language-Action Flow Model for General Robot Control
Related
Flow Matching · Action Expert · PaliGemma · π0.5 · π0-FAST · Physical Intelligence
Sources
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)
Physical-Intelligence/openpi GitHub 仓库 (Chinese)
As of
2025-09

See it in the full glossary →