Multimodal Large Language Model
多模态大语言模型MLLMCommonA large language model that can also understand non-text inputs like images, video, and audio.
A multimodal large language model uses an LLM as its ‘brain’ but can also take in non-text inputs such as images, video, and audio; GPT-4V brought wide attention to this direction in 2023. Yin and colleagues' survey breaks the typical architecture into three parts: a pretrained modality encoder (such as a CLIP or SigLIP vision encoder), a pretrained LLM, and a modality interface connecting the two — an MLP projector, a Q-Former, or cross-attention layers inserted into the LLM. Training usually proceeds in two stages: modality-alignment pretraining, then instruction tuning. An MLLM that only handles images and text is often called a vision-language model (VLM) instead. Its significance for embodied AI is that it brings internet-scale common sense and visual understanding: most VLAs are built by adding an action output on top of an MLLM or VLM, and embodied-reasoning models and task planners are also often fine-tuned from an MLLM.
ExampleLLaVA-1.5 uses a CLIP-ViT-L-336px vision encoder connected to the language model through an MLP projector, then is fine-tuned on visual instruction data; GPT-4o, Gemini, and Qwen2.5-VL are also multimodal large language models.
- Also called
- MLLM, Multimodal LLM
- Related
- Vision-Language Model · Large Language Model · Native Multimodal · Projector / Connector · Vision-Language-Action Model · Modality Alignment (Alignment Pretraining Stage)
- Sources
- A Survey on Multimodal Large Language Models (arXiv:2306.13549)
Improved Baselines with Visual Instruction Tuning (LLaVA-1.5, arXiv:2310.03744)