Embodied AI Glossary中文

OpenAI GPT Series

GPT 系列(GPT-4o / GPT-5)GPTCommon

OpenAI's large language model family; from GPT-4o onward it handles images and audio directly, often used by robots for high-level planning.

GPT stands for Generative Pre-trained Transformer, and here refers to the large language model family OpenAI has released since 2018: pretrained on massive text with next-token prediction, then fine-tuned and aligned. GPT-1 (2018, 117 million parameters), GPT-2 (2019, 1.5 billion), and GPT-3 (2020, 175 billion) successively showed that capability keeps improving with scale; GPT-4 (March 2023) began accepting image input; GPT-4o (May 2024, the “o” for “omni”) can process and generate text, images, and audio; GPT-5 launched on August 7, 2025, followed by iterations like 5.1 and 5.2; GPT-6 launched in September 2026, first opening an Astra version to paying users on September 4, then adding Sol and Luna versions on September 22. All of these are closed-source, accessible only through ChatGPT or the API. In embodied AI they're commonly used as an external “brain”: looking at images to understand a scene, breaking an instruction into steps, writing control code, or picking action primitives, then handing off to a low-level controller — the model itself doesn't output joint actions directly.

ExampleReKep first uses DINOv2 to mark numbered candidate keypoints on an image, then gives the annotated image and a language instruction to GPT-4o, which writes several Python constraint functions describing how the keypoints should relate to each other at each stage; an optimizer then solves for the end-effector's motion trajectory.

Also called
Generative Pre-trained Transformer, GPT-4o, GPT-5
Related
Large Language Model · Multimodal Large Language Model · OpenAI · Transformer · Next-Token Prediction · ReKep
Sources
Products and applications of OpenAI - Text generation(GPT-n 系列发布表,Wikipedia) (Chinese)
GPT-4o - Wikipedia
GPT-6 - Wikipedia
As of
2026-09

See it in the full glossary →