RECAP
CommonPhysical Intelligence's method for letting a VLA keep improving from its own experience plus human corrections.
RECAP is a training method Physical Intelligence released alongside π*0.6 in November 2025. Pure imitation learning can only ever match the level of its demonstrations, while a deployed robot generates a large stream of successes, failures, and human takeovers. RECAP combines all three kinds of data — offline demonstrations, trajectories from the robot's own autonomous runs, and corrections from an expert who teleoperates in to intervene during a run. It first trains a value function that predicts how many steps remain until task completion, uses that to judge whether each action was better or worse than average (its advantage), then trains the policy together with an “Advantage: positive / negative” text token as input. At inference, conditioning on “positive” pushes the model to output better actions. Deployment, retraining the value function, and fine-tuning the policy can be repeated in a loop. On tasks like folding laundry, assembling boxes, and making coffee, this more than doubled throughput and roughly halved the failure rate.
ExampleAfter training with RECAP, PI reports that π*0.6 can make espresso drinks continuously for 13 hours, and fold unfamiliar laundry in unfamiliar homes for more than two hours at a stretch.
- Also called
- RL with Experience and Corrections via Advantage-conditioned Policies
- Related
- π*0.6 · Advantage Conditioning · Value Function · Offline Reinforcement Learning · Human-in-the-Loop · Classifier-Free Guidance
- Sources
- π*0.6: a VLA That Learns From Experience (arXiv:2511.14759)
- As of
- 2025-11