Data Flywheel
数据飞轮EssentialA self-reinforcing loop: deployment produces data, the data improves the model, and the better model gets deployed even more widely.
A data flywheel is a self-reinforcing loop: a model is deployed, generates new data through real use — successes, failures, human corrections — and that data, once filtered and labeled, is used to improve the model. A better model can then take on more tasks and reach more deployments, which brings back still more data. NVIDIA defines it as a self-improving loop that keeps refining a model using data collected from its own interactions. The self-driving industry adopted this kind of approach early — Tesla calls its version the data engine. Robot data is expensive to collect, so companies generally want real-world deployment to help spread out that cost; but the flywheel can only start turning once a robot is already useful enough to be deployed in the real world.
ExamplePhysical Intelligence's π*0.6 uses a method called RECAP, which feeds a robot's own autonomous-execution data — from folding laundry, brewing espresso on a professional machine, and assembling boxes in real households — back into training, alongside expert teleoperation corrections. The paper reports throughput on the hardest tasks more than doubling, with the failure rate roughly cut in half.
- Also called
- Data Closed Loop, Data Engine
- Related
- Deployment Data Backflow · Human Intervention Data · RECAP · Self-improvement · Human-in-the-Loop · Fleet Learning (Learning While Deploying)
- Sources
- NVIDIA Glossary: What Is a Data Flywheel?
π*0.6: a VLA That Learns From Experience (arXiv 2511.14759) - As of
- 2025-11