Embodied AI Glossary中文

Native Multimodal

原生多模态Common

Training a model on multiple modalities together from the very start of pretraining, instead of bolting a vision module onto a language model.

Native multimodal means a model is trained from the start of pretraining on text, image, audio, and video data together, with every modality processed inside the same backbone network. It's the opposite of the ‘stitched-together’ approach, where a vision encoder and a language model are each trained separately first and then connected with a projector layer for further training. Google emphasized that Gemini was natively multimodal when it launched it in December 2023; Meta's Chameleon (2024) discretized images into tokens too and mixed them with text from the very beginning of training (early fusion); Llama 4 (April 2025) likewise claims to be natively multimodal and uses early fusion. Proponents argue this lets modalities combine more deeply and makes it easier to handle understanding and generation together; the cost is higher training expense and a harder balance to strike in the data mix across modalities. In embodied AI, unified multimodal models and world-action models that train video and action within the same sequence follow the same idea.

ExampleGemini was jointly pretrained on text, image, audio, and video data together from the start, whereas LLaVA attaches a vision encoder in front of an already-trained language model and then does further training — a stitched-together design.

Also called
Native Multimodal Model, Natively Multimodal
Related
Multimodal Large Language Model · Unified Multimodal Model · Multimodal Fusion · Visual Token · BAGEL · World Action Model
Sources
Google Blog: Introducing Gemini
Chameleon: Mixed-Modal Early-Fusion Foundation Models (arXiv:2405.09818)
Meta AI Blog: The Llama 4 herd
As of
2025-04

See it in the full glossary →