Embodied AI Glossary中文

Data Cleaning

数据清洗Common

Finding and fixing or removing bad data before training, such as failed trajectories, empty actions, or misaligned frames.

Data cleaning means identifying and fixing or removing corrupted, wrong, or irrelevant records in a dataset — a general step in data processing. Common problems in robot data include: trajectories that were never finished or where the operator made a mistake; idle frames with no motion at the start or end; all-zero actions; timestamps from multiple cameras and joint sensors that aren't aligned; dropped frames; sensor readings that jump unexpectedly; and language instructions that don't match what actually happened. This kind of noise teaches imitation learning to copy the pauses and jitter along with everything else. Cleaning methods include rule-based filtering, statistical anomaly detection, and manual spot-checking, or training a policy on a small cleaned batch first and seeing how it performs on the robot. Not all failure data should be deleted — failures and correction segments with the reason labeled are useful for learning to recover from mistakes.

ExampleThe OpenVLA paper attributes part of its improvement over RT-2-X to more careful data cleaning, such as removing all-zero actions from the Bridge data; AgiBot World's ablation study also shows that human-verified data raised the task-completion score by 0.18.

Also called
Data Cleansing
Related
Data Quality Control · Data Curation · No-op (Idle) Action Filtering · Valid (Usable) Data · Failure Data · Multi-sensor Time Synchronization / Timestamp Alignment
Sources
Wikipedia: Data cleansing
Kim et al. 2024: OpenVLA: An Open-Source Vision-Language-Action Model
AgiBot World Colosseo 技术报告 (arXiv 2503.06669) (Chinese)

See it in the full glossary →