Embodied AI Glossary中文

Autonomous Data Collection

自主数据采集Advanced

Letting a robot propose its own tasks, attempt them, and log the results, with little or no human teleoperation.

Autonomous data collection means a robot decides for itself what to do, executes it, and records the data with no human directing it step by step — a human only intervenes occasionally, mostly for safety oversight. Teleoperated collection needs one person per robot, so cost scales linearly with data volume; letting robots collect on their own means more robots simply means more data. Three difficulties stand out: the robot needs to propose meaningful and diverse tasks; it needs to automatically judge success or failure; and autonomously collected data tends to have a lower success rate and more variable quality, requiring methods that can learn from imperfect data. An early example is Google's 2016 Arm Farm, which used automatically judged grasp outcomes as labels; Google DeepMind's AutoRT uses a vision-language model to understand a scene and a large language model to propose tasks, orchestrating more than 20 robots; and Berkeley's SOAR uses a VLM to propose and judge tasks, autonomously collecting more than 30,000 trajectories across 5 tabletop environments to improve a policy.

ExampleAutoRT ran more than 20 robots across multiple office buildings, with a large language model proposing tasks, collecting about 77,000 real-robot trajectories total through a mix of teleoperation and autonomous execution.

Also called
Robot Self-Collection
Related
Google Arm Farm · AutoRT · Data Flywheel · Self-improvement · Real-World Reinforcement Learning · Success Detector
Sources
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents (arXiv 2401.12963)
Autonomous Improvement of Instruction Following Skills via Foundation Models (SOAR, arXiv 2407.20635)
Deep Learning for Robots: Learning from Large-Scale Interaction (Google Research Blog)

See it in the full glossary →