Embodied AI Glossary中文

Google Gemini

Gemini 系列(谷歌多模态大模型)Common

Google DeepMind's family of natively multimodal large models, also the foundation for robotics models such as Gemini Robotics.

Gemini is Google DeepMind's family of multimodal large models, succeeding LaMDA and PaLM 2. The first version, Gemini 1.0, launched in December 2023 in three sizes — Ultra, Pro, and Nano. From the start of training it handles text, images, audio, video, and code together, rather than training a language model first and bolting on a vision module afterward — an approach often called “natively multimodal.” Iteration has been fast: 1.5 in 2024 (very long context), 2.0 and 2.5 from late 2024 into 2025, Gemini 3 in November 2025, and as of September 2026 several versions within the 3.x series. In embodied AI it's used two ways: directly, to look at images, understand a scene, break down tasks, and do high-level planning; and as a backbone for training robot models — Google's Gemini Robotics (a VLA) and Gemini Robotics-ER (embodied reasoning), both released in March 2025, are built on top of Gemini 2.0.

ExampleGemini Robotics adds “physical action” as a new output modality on top of Gemini 2.0: after seeing the camera feed and hearing an instruction, the model directly outputs actions that control the robot.

Also called
Gemini
Related
Multimodal Large Language Model · Native Multimodal · Gemini Robotics · Gemini Robotics-ER · Google DeepMind · Vision-Language Model
Sources
Gemini: A Family of Highly Capable Multimodal Models (arXiv 2312.11805)
Gemini (language model) - Wikipedia
Gemini Robotics brings AI into the physical world (Google DeepMind blog)
As of
2026-09

See it in the full glossary →