Embodied AI Glossary中文

Data Mixture

数据配比Common

The proportion each data source is given when training on a mix of multiple sources at once.

Data mixture refers to how often each source gets sampled when training on a combination of data from different origins — different robots, different tasks, simulation versus real data, web text and images, human video. Source sizes can differ enormously; mixing purely in proportion to raw counts lets big datasets drown out small ones, while poorly tuned weights instead make the model lopsided. Common approaches include hand-tuned weights, such as the recipe Octo and OpenVLA use on Open X-Embodiment (informally called the Magic Soup); π0's rule of weighting each task-robot combination by n^0.43 (n being that combination's sample count), which down-weights combinations with too much data; and Stanford's Re-Mix, which learns weights automatically with distributionally robust optimization, reporting a 38% average improvement over uniform weighting. Mixtures are also commonly changed across training stages, leaning toward diversity during pretraining and toward quality during post-training.

ExampleIn π0's pretraining data, 9.1% comes from open datasets such as OXE, Bridge v2, and DROID, with the rest from Physical Intelligence's own collection, and each task-robot combination is then weighted by the n^0.43 rule.

Also called
Data Recipe, Data Mixture Weights
Related
OXE Magic Soup · Co-training · Heterogeneous Data · Cross-Embodiment Data · Data Curation · Open X-Embodiment
Sources
Re-Mix: Optimizing Data Mixtures for Large Scale Imitation Learning (arXiv 2408.14037)
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)
As of
2024-10

See it in the full glossary →