Embodied AI Glossary中文

Mixture-of-Transformers

混合 Transformer 架构MoTAdvanced

Giving each modality its own set of Transformer parameters, while every layer's self-attention still lets them all see each other.

Mixture-of-Transformers was proposed in November 2024 by Stanford's Weixin Liang and researchers at Meta. It splits the feedforward network, attention projection matrices, and layer normalization parameters by modality: text tokens use one set of weights, image tokens use another, but each layer's self-attention is still computed over the whole sequence, so the modalities interact as usual. Unlike a Mixture of Experts (MoE, where a router picks experts per token), MoT assigns a fixed division of labor by modality and needs no routing to be learned. The paper reports that at 7B scale, on a combined text-and-image generation setup, it matches a dense model's performance using only about 55.8% of the compute. In embodied AI, π0's 'VLM backbone + action expert' design is a similar approach (the paper describes it as resembling an MoE with two members), while ByteDance's BAGEL adopts the MoT structure directly.

Exampleπ0's roughly 3-billion-parameter backbone comes from PaliGemma, plus a separate action expert of about 300 million parameters: image and text tokens use the backbone's weights, and robot state and noisy action tokens use the action expert's weights, with the two parts visible to each other in every layer's attention.

Also called
MoT, Mixture-of-Transformer-Experts
Related
Mixture of Experts · Action Expert · π0 · BAGEL · Unified Multimodal Model · Self-Attention
Sources
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models (arXiv:2411.04996)
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)
Emerging Properties in Unified Multimodal Pretraining (BAGEL, arXiv:2505.14683)
As of
2025-05

See it in the full glossary →