Generalist Policy
通用策略(通才策略)EssentialA single control policy that can perform many tasks, and often work across many settings or robots.
A policy is a model that maps observations to actions. A generalist policy is a single policy trained on large-scale, multi-task — often cross-robot — data, able to follow a language instruction or a goal image to complete many different tasks, and to generalize somewhat to new settings. The contrast is a specialist policy, trained for just one task or one robot. Generalist policies are usually pretrained at scale first, then fine-tuned to a specific robot and task with a small amount of target-domain data, an approach modeled on large language models. Notable examples include the open-source Octo (2024, trained on 800,000 trajectories from Open X-Embodiment and fine-tunable to a new robot in a few hours on a consumer GPU), OpenVLA, and Physical Intelligence's π0 series. Most vision-language-action (VLA) models today aim to be generalist policies.
ExampleGive Octo either a spoken instruction or an image of the completed task, and it can output robot-arm actions accordingly, without needing a separate model trained for each task.
- Also called
- Generalist Robot Policy
- Related
- Policy · Specialist Policy · Octo · π0 · Vision-Language-Action Model · Cross-Embodiment
- Sources
- Octo: An Open-Source Generalist Robot Policy
π0: A Vision-Language-Action Flow Model for General Robot Control - As of
- 2024-10