Embodied AI Glossary中文

Hindsight Relabeling

事后重标注Advanced

After data is collected, relabeling a trajectory's goal or instruction to match whatever it actually achieved.

The idea behind hindsight relabeling is: a trajectory that failed to reach its original goal still ended up somewhere, so that actual outcome can be relabeled as the goal after the fact — turning a failed sample into a success at “achieving a different goal.” The technique became popular through OpenAI's 2017 Hindsight Experience Replay (HER), which addresses how little a policy can learn under sparse rewards (a reward given only on success), and was validated on robot-arm tasks like pushing, sliding, and pick-and-place. It was later extended to imitation learning and language-conditioned policies: a large amount of manipulation data is collected with no task label at all, and afterward it's labeled with an image goal or a language instruction based on what the footage shows. Google's DIAL, for instance, uses a vision-language model like CLIP to automatically add language instructions to 80,000 demonstrations, 96.5% of which had no human annotation to begin with. The precondition for this technique is that the policy is conditioned on a goal or instruction — otherwise there's nothing to relabel.

ExampleA robot arm meant to push a block to point A instead pushes it to point B; after relabeling, this trajectory is stored as a successful demonstration with “goal = point B,” and used to train a goal-conditioned policy.

Also called
Hindsight Goal Relabeling
Related
Hindsight Experience Replay · Goal-Conditioned Reinforcement Learning · Sparse Reward · Language Annotation · Auto-labeling · Instruction Augmentation
Sources
Hindsight Experience Replay (arXiv)
Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models (DIAL, arXiv)

See it in the full glossary →