Embodied AI Glossary中文

π0.5

Essential

PI's 2025 VLA that handles long-horizon tasks like tidying a kitchen in real homes it has never seen.

π0.5 is the upgraded version of π0 that Physical Intelligence released on April 22, 2025, focused on open-world generalization — getting a robot to work in homes it never visited during training. The approach co-trains on a mix of data sources: roughly 400 hours of mobile-manipulator data collected across about 100 homes makes up only about 2.4% of the pretraining examples, with the rest coming from other robots, lab data, web image-text question answering, and high-level semantic annotations. At inference, the model first predicts the next subtask in words (such as “pick up the plate”), then generates low-level actions conditioned on that — similar to chain-of-thought. During training, the pretraining stage compresses actions into discrete FAST tokens for efficiency, and post-training attaches a flow-matching action expert that outputs continuous actions. The paper tested it in 3 real homes excluded from training, where it completed multistep tasks like tidying a kitchen or straightening a bedroom. The weights were open-sourced in openpi in September 2025.

ExampleGiven the instruction “clean the kitchen” in a home it has never visited, π0.5 first states a subtask like “put the plate in the sink,” then executes the corresponding actions, working through the task step by step.

Also called
pi0.5, pi05, π0.5: a Vision-Language-Action Model with Open-World Generalization
Related
π0 · Co-training · π0-FAST · Open-world · Long-horizon Task · π*0.6
Sources
π0.5: a Vision-Language-Action Model with Open-World Generalization (arXiv 2504.16054)
Physical-Intelligence/openpi GitHub 仓库 (Chinese)
As of
2025-09

See it in the full glossary →