Feature-wise Linear Modulation
FiLM 特征调制FiLMAdvancedUsing conditioning information to compute a per-channel scale and shift that modulates a network's intermediate features.
FiLM was proposed by Ethan Perez, Aaron Courville, and colleagues (AAAI 2018) as a general-purpose layer for injecting conditioning information into a neural network. The idea is simple: a small network computes a scale γ and a shift β for each feature channel from the condition (such as a language-instruction vector), and the intermediate feature is transformed into γ·x + β. The original paper roughly halved the best error rate at the time on the CLEVR visual-reasoning benchmark. It's common in robotics: RT-1 uses FiLM to inject the language instruction into its pretrained EfficientNet image encoder, zero-initializing the layers that produce γ and β so FiLM starts out as an identity transform and doesn't disturb the pretrained weights; the CNN version of Diffusion Policy also uses FiLM at every convolutional layer to inject observation features.
ExampleIn RT-1, the instruction 'pick up the coke can' is first turned into a vector by the Universal Sentence Encoder, then modulates EfficientNet's layer features through FiLM, so the same image produces different visual features under different instructions.
- Also called
- FiLM, FiLM Layer, FiLM Conditioning
- Related
- Adaptive Layer Normalization · Language-conditioned Policy · RT-1 · Diffusion Policy · EfficientNet · Cross-Attention
- Sources
- FiLM: Visual Reasoning with a General Conditioning Layer (Perez et al., AAAI 2018)
RT-1: Robotics Transformer for Real-World Control at Scale (arXiv 2212.06817)
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)