Data Diversity
数据多样性CommonHow broadly training data varies across scenes, objects, tasks, and viewpoints.
Data diversity describes how much a dataset varies along dimensions such as environment, object, lighting, camera viewpoint, task, operator, and robot embodiment — a different question from “how many demonstrations there are.” A robot policy easily memorizes the specific table and objects it trained on and fails in a new room, so insufficient diversity is one of the main reasons generalization is poor. A 2024 data-scaling-law study by Tsinghua's Yang Gao and colleagues found that a policy's performance in new environments and on new objects scales mainly with the number of distinct environments and objects, and that collecting more demonstrations in the same environment quickly stops helping. Physical Intelligence's π0.5 trained on mobile-manipulation data from roughly 100 real homes and was still able to do tidying-type tasks in homes it had never seen. Deliberately varying the scene, the objects, and the person collecting, during data collection, is what raises diversity.
ExampleThe π0.5 paper's experiments show that the more homes the training data came from, the better the policy performed in test homes; training on data from 104 locations came close to matching a control model trained directly on the test home's own data.
- Also called
- Data Coverage
- Related
- Generalization · Scene Generalization · Object Generalization · Data Scaling Laws in Imitation Learning (Robotic Manipulation) · In-the-wild Data · Data Mixture
- Sources
- Data Scaling Laws in Imitation Learning for Robotic Manipulation (arXiv 2410.18647)
π0.5: a Vision-Language-Action Model with Open-World Generalization (arXiv 2504.16054) - As of
- 2025-04