Suboptimal (Noisy) Demonstrations
次优演示AdvancedHuman demonstration data that includes wasted motion, hesitation, mistakes, or inconsistent quality across operators.
Suboptimal demonstrations are demonstrations that fall short of ideal: slow motion, pauses, back-and-forth trial and error, mistakes followed by recovery, or inconsistency because different operators have different skill levels and habits. Almost all real-world collected data has some of these issues. Behavioral cloning (directly imitating the demonstrated actions) learns from the good and bad actions alike, so the more mixed the data, the more hesitant and unstable the resulting policy tends to be. Common countermeasures include: filtering and quality-checking the data after collection; weighting samples by reward or advantage (how much better an action is than average), as in advantage-weighted regression; using offline reinforcement learning to extract better behavior from mixed-quality data; or feeding quality information in as a conditioning signal, so that only the “good” behavior is requested at inference time, as in RECAP's advantage conditioning.
Examplerobomimic's Multi-Human dataset was collected by 6 operators of varying skill (2 each rated “worse,” “okay,” and “better”), each recording 50 successful trajectories for 300 total, specifically to study how mixed-quality data affects imitation learning.
- Also called
- Imperfect Demonstrations, Mixed-Quality Demonstrations
- Related
- Behavior Cloning · Data Curation · Data Quality Control · Advantage-Weighted Regression · Offline Reinforcement Learning · RECAP
- Sources
- robomimic v0.1 Datasets (PH / MH / MG)
What Matters in Learning from Offline Human Demonstrations for Robot Manipulation (arXiv 2108.03298)