Instruction Augmentation
指令增强AdvancedAutomatically writing or rewriting language instructions for robot trajectories, so the same motion gets many phrasings.
Instruction augmentation is a family of data-processing methods for language-conditioned policies (policies that act based on text instructions): rather than collecting new robot motion, it adds or rewrites the language label on trajectories that already exist. Writing an instruction by hand for every demonstration is expensive, and the wording tends to be narrow, so a model can fail to understand the same task described a different way. A representative example is DIAL from the Google robotics team, proposed in 2022 and published at RSS 2023: it first fine-tunes CLIP (an image-text matching model) on a small amount of human-labeled data, then has a large language model propose candidate instructions, uses CLIP to score and pick the ones that match what the trajectory's footage actually shows, and relabels 80,000 demonstrations this way (96.5% of which had no human language label at all), before running behavioral cloning. A simpler version just has a large model rewrite the original instruction into several synonymous phrasings. It follows the same underlying idea as hindsight relabeling and automatic annotation.
ExampleThe policy trained with DIAL was tested on new instructions that never appeared in the original 60-instruction dataset, spanning three categories: spatial-relationship descriptions, differently worded descriptions, and entirely new semantic skills.
- Also called
- Language Augmentation
- Related
- Language Annotation · Hindsight Relabeling · Auto-labeling · Data Augmentation · Language-conditioned Policy · CLIP
- Sources
- Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models (arXiv 2211.11736)
DIAL 项目主页 (Chinese)