Multimodal Fusion
多模态融合CommonCombining information from different sources — image, language, robot state — into a form a model can use jointly.
Multimodal fusion means combining information from different modalities — image, text, depth, touch, joint state, and so on — into a representation the model can use jointly. Baltrušaitis and colleagues' 2017 survey lists it as one of five core challenges in multimodal machine learning, alongside representation, translation, alignment, and co-learning. Fusion approaches are usually grouped by where the combination happens: early fusion concatenates features or tokens from each modality right at the start and processes them together; mid fusion encodes each modality separately first, then merges them partway through the network, often via cross-attention; and late fusion lets each modality produce its own result and combines them only at the final decision. Embodied AI is inherently multimodal — a VLA has to look at camera images, read a language instruction, and sense its own state at the same time — and the fusion approach directly determines whether the model can map an instruction like ‘pick up the red cup’ onto the right object and action.
ExampleRT-1 uses FiLM layers to inject the language instruction's embedding into its EfficientNet-B3 image encoder, so visual features carry task information from early in the network; π0, by contrast, puts image, language, state, and action tokens into the same Transformer and lets attention handle the fusion.
- Also called
- Early / Late Fusion
- Related
- Multimodal Large Language Model · Cross-Attention · Feature-wise Linear Modulation · Projector / Connector · Visuo-Tactile Fusion · Multi-Sensor Fusion
- Sources
- Multimodal Machine Learning: A Survey and Taxonomy (arXiv:1705.09406)
Wikipedia: Multimodal learning
RT-1: Robotics Transformer for Real-World Control at Scale (arXiv:2212.06817)