Embodied AI Glossary中文

Data Annotation

数据标注Common

Adding text descriptions, segment boundaries, bounding boxes, and other labels to raw data to tell a model what it is.

Data annotation means attaching labels to raw data that a human can understand and a model can use as a supervision signal — the classic example is labeling an image's category or drawing a bounding box. In robot data, actions are already recorded automatically during collection, so what mostly needs adding is: the language instruction for the whole task, where a task splits into subtasks and their start/end times, bounding boxes around target objects, and whether the attempt succeeded or failed and why. These labels determine whether a language-conditioned policy can actually understand instructions, and whether a long-horizon task can be learned as a sequence of steps. Annotation can be done entirely by hand, or generated automatically by a vision-language model and then checked by a human. Labeled data costs far more than raw data, and inconsistency between annotators can directly hurt model performance.

ExampleThe AgiBot World 2026 dataset provides three layers of annotation: task-frame segmentation with subtask instructions, 2D bounding boxes for object interactions, and step-level instruction segmentation with atomic skills; DROID's December 2024 update added 3 natural-language descriptions to each of about 75,000 successful trajectories.

Also called
Data Labeling
Related
Language Annotation · Subtask Segmentation · Auto-labeling · Action Label · Data Quality Control · Hindsight Relabeling
Sources
Wikipedia: Labeled data
Hugging Face: agibot-world/AgiBotWorld2026
DROID 数据集项目页 (Chinese)

See it in the full glossary →