Subtask Segmentation
子任务切分AdvancedSplitting one long demonstration into steps, marking each segment's start and end frame with a description.
Subtask segmentation is a data-annotation step: a complete long-horizon demonstration (such as “clear the table”) is split by meaning into subtask segments (“pick up the plate,” “put it in the sink”), with the start and end frame of each recorded, usually along with a language description and a skill type. It lets long-horizon task data be used step by step: when training a hierarchical policy, the high-level model learns to predict which subtask comes next, while the low-level model learns to execute a single subtask; segmentation can also be used to compute per-skill success rates, build a progress-based reward, or discard failed segments. Segmentation methods include manual annotation, rule-based automatic segmentation (for example, at points where the gripper opens or closes, or velocity drops near zero), and automatically finding boundaries using visual representations or a VLM — UVD, for instance, discovers subgoals by detecting sudden shifts in a pretrained visual representation.
ExampleThe AgiBot World 2026 dataset provides step-level instruction_segments annotations for every trajectory, with each segment recording a skill type (such as Pick), one instruction sentence, and its start/end frames — for example, “left arm picks up the red-capped drink from the shopping cart” corresponds to frames 284–493.
- Also called
- Trajectory Segmentation, Skill Segmentation
- Related
- Long-horizon Task · Language Annotation · Auto-labeling · Skill Primitive · Hierarchical Architecture · AgiBot World
- Sources
- Hugging Face: agibot-world/AgiBotWorld2026 数据集卡片 (Chinese)
Universal Visual Decomposer (arXiv 2310.08581)