Language Annotation
语言标注CommonAttaching a natural-language description to a robot trajectory, such as 'put the red cup in the sink,' so a model can follow instructions.
Language annotation means attaching a text description to robot data, most commonly one task instruction per trajectory, though it can go finer, with one sentence per subtask. It is what makes it possible to train language-conditioned policies and VLA models that act on instructions. Sources fall roughly into three groups: deciding the instruction before collection and having the operator follow it; writing it in afterward by a human watching the video, called hindsight relabeling; and generating or rewriting it automatically with a vision-language model. Annotation quality directly affects how well a model understands instructions — if the same motion is only ever described one way, the model tends to latch onto that fixed phrasing. Large datasets therefore commonly pair one trajectory with several different phrasings, and split long tasks into text-labeled subtask segments to help hierarchical models learn.
ExampleGoogle's Language-Table used hindsight relabeling to produce nearly 600,000 language-labeled tabletop pushing trajectories; DROID's December 2024 update added 3 language descriptions to each of about 75,000 successful trajectories.
- Also called
- Instruction Annotation
- Related
- Data Annotation · Hindsight Relabeling · Subtask Segmentation · Auto-labeling · Language-conditioned Policy · Instruction Augmentation
- Sources
- Interactive Language / Language-Table 项目主页 (Chinese)
DROID 数据集主页 (Chinese) - As of
- 2024-12