Embodied AI Glossary中文

π*0.6

Common

π0.6 after reinforcement learning with RECAP, able to keep improving from demonstrations, corrections, and its own experience.

π*0.6 is a vision-language-action model Physical Intelligence released in November 2025. Its base, π0.6, builds on π0.5 but switches to a Gemma 3 4B vision-language backbone and grows the action expert to 860 million parameters; the asterisk marks that it was further trained with reinforcement learning using the RECAP method. A model trained purely on human demonstrations tends to drift further off course after a small mistake (compounding error), making it hard to succeed reliably. RECAP trains on three kinds of data at once: human demonstrations, corrections made when an expert teleoperates in to take over, and the robot's own successes and failures from autonomous runs. It first trains a value function that estimates how many steps remain to completion, uses that to judge whether a stretch of actions made things better or worse (its advantage), then trains the model together with an “Advantage: positive / negative” text condition, and at deployment only asks it for the “positive” behavior. On the hardest tasks, this more than doubled throughput and roughly halved the failure rate.

ExampleAfter RECAP training, π*0.6 made espresso drinks continuously for 13 hours, folded unfamiliar laundry without interruption for more than two hours in a new home, and assembled real shipping boxes on a factory floor.

Also called
π0.6, pi0.6, pi*0.6, pi-star-0.6
Related
RECAP · Vision-Language-Action Model · Reinforcement Fine-Tuning (RL Fine-Tuning) · Advantage Conditioning · Value Function · Physical Intelligence
Sources
π*0.6: a VLA That Learns From Experience (arXiv 2511.14759)
π*0.6: a VLA that Learns from Experience (Physical Intelligence blog)
As of
2025-11

See it in the full glossary →