Embodied AI Glossary中文

Specialist Policy

专用策略Common

A policy trained only for a single robot, a single task, or a single setting — the counterpart to a generalist policy.

A specialist policy is a control policy trained from scratch on data collected specifically for one robot, one task, or even one particular environment. This used to be how robot learning was mostly done: switch to a different robot or task, and you collect new data and train a new model from scratch. It tends to perform well on its own task and needs relatively little data, but it transfers poorly. The Open X-Embodiment comparison (2023) showed that in low-data settings, RT-1-X, trained on mixed data from 22 robots, outperformed the original single-dataset methods by about 50% on average. The common approach now is to pretrain a generalist policy first, then fine-tune it into a specialist policy with a small amount of task-specific data. Note that “expert policy” in imitation learning has a different meaning, referring to the expert that provides the demonstrations.

ExampleTraining an ACT policy from scratch using only “unscrew the bottle cap” demonstrations collected on one particular ALOHA robot: it works on that robot for that task, but switching to folding clothes or to a different arm requires collecting new data and retraining.

Also called
Single-task Policy
Related
Generalist Policy · Policy · Fine-tuning · Cross-Embodiment · RT-X · Open X-Embodiment
Sources
Open X-Embodiment: Robotic Learning Datasets and RT-X Models (project page)
Open X-Embodiment: Robotic Learning Datasets and RT-X Models (arXiv)
Octo: An Open-Source Generalist Robot Policy
As of
2023-10

See it in the full glossary →