Embodied AI Glossary中文

Retrieval-Based Data Selection

数据检索Advanced

Using a handful of demonstrations for a target task to search a large dataset for similar clips, then training on the combined set.

Retrieval-based data selection is a family of methods for choosing training data: collect a few demonstrations for the target task, then search a large existing robot dataset for trajectories or segments that are visually or behaviorally similar, and train on the retrieved data together with the target demonstrations. It addresses a tradeoff: training on all available data mixed together risks negative transfer, where unrelated data interferes with learning, while training on just a handful of demonstrations isn't enough data on its own. Representative work includes Stanford's Behavior Retrieval (2023); the University of Washington's STRAP (ICLR 2025), which retrieves at the sub-trajectory level using vision-foundation-model features and dynamic time warping for matching; and Stanford's IWR (CoRL 2025), which uses importance weighting to correct for retrieval bias. It addresses the same underlying problem as data filtering and data mixing ratios.

ExampleTo teach a robot to pick up a cup, put it in a drawer, and close the drawer, only a few demonstrations were collected; STRAP retrieves sub-segments involving “picking up a cup” and “opening/closing a drawer” from a large offline dataset and trains the policy on those together with the target demonstrations.

Also called
Data Retrieval
Related
Data Curation · Data Mixture · Few-shot · Positive / Negative Transfer · Open X-Embodiment · Behavior Cloning
Sources
Behavior Retrieval: Few-Shot Imitation Learning by Querying Unlabeled Datasets (arXiv)
STRAP: Robot Sub-Trajectory Retrieval for Augmented Policy Learning (arXiv)
Data Retrieval with Importance Weights for Few-Shot Imitation Learning (arXiv)
As of
2025-09

See it in the full glossary →