LLaVA
LLaVA(视觉指令微调架构)CommonA 2023 open-source multimodal model that connects a vision encoder to an LLM through a projection layer, then fine-tunes on instructions.
LLaVA was proposed in April 2023 by Haotian Liu, Chunyuan Li, and colleagues at the University of Wisconsin-Madison, Microsoft Research, and Columbia University; the paper, “Visual Instruction Tuning,” was an oral presentation at NeurIPS 2023. The architecture is simple: a CLIP ViT-L/14 vision encoder extracts image features, which pass through a projection layer (mapping visual features into the language model's word-embedding space) and into a Vicuna large language model. Training has two stages: first the language model is frozen and only the projection layer is trained, to align the two modalities; then the whole thing is fine-tuned end to end. The key ingredient is data: GPT-4, given only text, generated 158,000 image-instruction pairs covering conversations, detailed descriptions, and complex reasoning. LLaVA-1.5, from October 2023, replaced the projection layer with a two-layer MLP, trained in about a day on a single 8-GPU A100 node using only public data, and set the state of the art across 11 benchmarks. “Vision encoder + projection layer + LLM + instruction tuning” became the common recipe for most later open-source VLMs, and many VLAs add an action output on top of this kind of VLM.
ExampleWhen generating training data, GPT-4 never sees the image itself — only a text description of it and the object bounding-box coordinates — and from that it writes multi-turn questions, answers, and reasoning problems about the image; those Q&A pairs are then paired with the actual image to train LLaVA.
- Also called
- Large Language and Vision Assistant, Visual Instruction Tuning, LLaVA-1.5
- Related
- Vision-Language Model · Multimodal Large Language Model · Instruction Tuning · Projector / Connector · CLIP · Prismatic VLMs
- Sources
- Visual Instruction Tuning (arXiv 2304.08485)
Improved Baselines with Visual Instruction Tuning (LLaVA-1.5, arXiv 2310.03744)
LLaVA 项目主页 (Chinese) - As of
- 2023-10