Embodied AI Glossary中文

EmbodiedGPT

Advanced

A 2023 embodied multimodal model from HKU and Shanghai AI Lab that uses chain of thought to generate step-by-step plans.

EmbodiedGPT was released in May 2023 by the University of Hong Kong, Shanghai AI Lab, and others, published at NeurIPS 2023, and is one of the early landmark works applying large models to embodied planning. The authors selected clips from Ego4D first-person video and annotated them, in chain-of-thought form, as step-by-step sub-goals, building a planning dataset called EgoCOT; they then used prefix tuning to adapt a 7B language model to this kind of data, so it outputs a step-by-step plan after seeing an image. The key design is treating the plan the language model generates as a query, used to extract task-relevant features from the image, which are handed to a low-level policy network — closing the loop between high-level planning and low-level control.

ExampleOn the Franka Kitchen and Meta-World simulated control tasks, EmbodiedGPT's success rate is 1.6 times and 1.3 times that of a BLIP-2 baseline fine-tuned on Ego4D, respectively.

Also called
Vision-Language Pre-Training via Embodied Chain of Thought
Related
Embodied Chain-of-Thought · Chain-of-Thought · Egocentric Video · Ego4D · LLM-based Task Planning · Franka Kitchen
Sources
EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought (arXiv 2305.15021)
NeurIPS 2023 论文页 (Chinese)
As of
2023-12

See it in the full glossary →