Embodied Foundation Model
具身大模型EssentialA large model pretrained on data spanning many robots and tasks, adaptable to many different robots and jobs.
Embodied foundation model extends the foundation-model idea into robotics. A foundation model is one pretrained on large-scale, diverse data, usually with self-supervised learning, meaning it builds training targets from the data itself rather than from human labels, and can then adapt to many downstream tasks, GPT being an example. Embodied foundation models likewise pretrain first on data spanning many robots and tasks, then post-train on a small amount of data for a specific task or a specific robot body, rather than training a dedicated policy from scratch for every task. The term has no fixed boundary: narrowly it often means a generalist policy that outputs actions directly, such as π0, GR00T N1, or Gemini Robotics; broadly it also includes embodied reasoning models responsible for understanding and planning, and world models used to generate data or evaluate policies. Firoozi and colleagues' 2023 survey identifies the main obstacles as scarce robot data, missing safety guarantees, and real-time requirements.
Exampleπ0 was pretrained on more than 10,000 hours of robot data, including self-collected data spanning 7 robot configurations and 68 tasks, and then post-trained on curated data for long-horizon tasks such as folding laundry and assembling a box.
- Also called
- Robot Foundation Model, RFM
- Related
- Foundation Model · Vision-Language-Action Model · Generalist Policy · Embodied Reasoning Model · World Model · Pre-training
- Sources
- Foundation Models in Robotics: Applications, Challenges, and the Future (arXiv 2312.07843)
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)