Modality Alignment (Alignment Pretraining Stage)
模态对齐CommonMapping features from a new modality, like images, into a representation space the language model can already understand.
Modality alignment means bringing representations of different modalities — image, language, robot state, action — into a space where they correspond to one another. Narrowly, it often refers to the first stage of training a multimodal large model, “alignment pretraining”: LLaVA (2023), for example, freezes both a CLIP vision encoder and a large language model and trains only the projection layer in between, using about 595,000 image-text pairs, so image features turn into visual tokens the language model can process; the second stage then fine-tunes the projection layer together with the language model on instruction data. Aligning first, on its own, lets the randomly initialized projection layer learn to produce reasonable visual tokens before the language model itself is unfrozen. More broadly, CLIP's use of contrastive learning to pull image and text embeddings into the same space also counts as modality alignment. VLA models face the same problem when attaching a state encoder and an action head to a vision-language model.
ExampleLLaVA's first stage trains only the projection matrix for 1 epoch, at a learning rate of 2e-3, on a filtered subset of CC3M (595,000 image-text pairs); its second stage freezes the vision encoder and fine-tunes the projection layer and language model together on 158,000 multimodal instruction examples.
- Also called
- Alignment Pretraining, Feature Alignment Pretraining, Cross-modal Alignment
- Related
- Projector / Connector · Vision-Language Model · Multimodal Large Language Model · CLIP · Backbone Freezing · Knowledge Insulation
- Sources
- Visual Instruction Tuning (LLaVA, arXiv 2304.08485)