Embodied AI Glossary中文

Deployment Data Backflow

数据回流Advanced

A term from China's robotics industry: feeding data generated during real-world robot deployment back into training, then redeploying the updated model.

Data backflow is a common term in China's embodied-AI industry for a loop: data generated while a robot works in the real world — autonomously executed trajectories, failure cases, and clips where a human took over and corrected it — is logged, sent back, cleaned, and annotated, then added to training; the updated model is redeployed, and the cycle repeats. It's the key link that gets a “data flywheel” actually spinning: teleoperation data collected in a training facility covers a limited distribution, and errors that show up in the field are the best way to expose a model's real weaknesses. CAICT's Embodied Intelligence Development Report (2025) argues that robots need to move past pure “demonstration” data as quickly as possible and establish a continuous data-backflow loop from real deployment. The hard part is cost: if every robot needs an operator watching and ready to take over the whole time, it's difficult to make commercially viable.

ExampleWhen training π*0.6, Physical Intelligence used the RECAP method to keep training on a mix of data autonomously generated while robots made espresso, folded laundry, and assembled boxes, together with human correction data; the company reported that throughput on some of the hardest tasks more than doubled, while the failure rate dropped by roughly half.

Also called
Data Backflow, Field Data Feedback Loop
Related
Data Flywheel · Human Intervention Data · Recovery and Correction Data · Fleet Learning (Learning While Deploying) · RECAP · π*0.6
Sources
中国信通院《具身智能发展报告(2025年)》 (Chinese)
具身智能迈向2.0:数据采集从训练场走向真实世界(科学网转澎湃新闻) (Chinese)
π*0.6: a VLA That Learns From Experience (arXiv)
As of
2026-09

See it in the full glossary →