Adaptive Layer Normalization
自适应层归一化AdaLNAdvancedComputing layer normalization's scale and shift dynamically from conditioning information, such as a diffusion timestep, instead of fixing them.
Layer normalization standardizes each token's features, then multiplies by a scale γ and adds a shift β; ordinarily these two parameters are fixed after training. Adaptive layer normalization instead has a small MLP regress γ and β from a conditioning vector — a diffusion timestep, a class label, a language feature — so that whenever the condition changes, the distribution of the whole layer's features changes with it, 'injecting' the condition into the network; the idea is the same as FiLM. It builds on the adaptive normalization used in GANs and U-Net diffusion models. Peebles and Xie used it systematically in the diffusion Transformer (DiT) in 2022 and introduced adaLN-Zero: an extra gating coefficient is regressed and initialized to zero, so each Transformer block starts out as an identity mapping, which makes training more stable. In embodied AI, NVIDIA's GR00T N1 uses AdaLN in its DiT action module to inject the denoising step, while using cross-attention to receive VLM features.
ExampleThe openpi implementation of π0.5 encodes the flow-matching timestep with a two-layer MLP into 'adarms_cond,' which modulates the normalization layers on the action tokens using adaptive RMSNorm.
- Also called
- AdaLN, adaLN-Zero, adaRMS
- Related
- Feature-wise Linear Modulation · Diffusion Transformer · Normalization Layers · Action Expert · NVIDIA Isaac GR00T N1 · Cross-Attention
- Sources
- Scalable Diffusion Models with Transformers (DiT, arXiv 2212.09748)
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)
openpi pi0.py 源码 (Chinese)