Embodied AI Glossary中文

Data Curation

数据筛选Common

Picking out the subset of a large robot dataset that actually helps training, and discarding the harmful part.

Data curation means selecting and weighting a dataset before training: dropping demonstrations that are error-prone, hesitant, or inconsistent in technique, and keeping the high-quality, broadly covering portion. It differs from data cleaning (fixing formats, removing bad frames) in that it asks which data will actually make the policy better. Robot demonstrations are often collected by many people and vary widely in quality, and imitation learning will copy bad habits right along with good ones, so more data is not automatically better. Notable methods include DemInf, proposed by researchers at Stanford and Google DeepMind in 2025, which scores each demonstration using the mutual information between state and action, and CUPID, from CoRL 2025, which uses influence functions to estimate each demonstration's contribution to policy success rate — the paper reports that using less than 33% of the curated data trained a then-state-of-the-art diffusion policy on the RoboMimic benchmark.

ExampleIn the DemInf paper, on real ALOHA and Franka setups, demonstrations collected by multiple people were scored one by one, the lowest-scoring batch was removed, and the policy trained on the remainder outperformed one trained on the full dataset.

Also called
Data Filtering
Related
Data Cleaning · Data Quality Control · Data Mixture · Valid (Usable) Data · Suboptimal (Noisy) Demonstrations · Imitation Learning
Sources
Robot Data Curation with Mutual Information Estimators (arXiv 2502.08623)
CUPID: Curating Data your Robot Loves with Influence Functions (arXiv 2506.19121)
As of
2025-09

See it in the full glossary →