Masked Autoencoder
掩码自编码器MAEAdvancedHiding most patches of an image and training a model to reconstruct them, as self-supervised visual pretraining.
The masked autoencoder is a self-supervised visual-pretraining method proposed by Kaiming He and colleagues at Meta FAIR in 2021. An image is split into patches, about 75% of them are randomly masked, the encoder processes only the visible patches, and a lightweight decoder reconstructs the masked-out pixels from that encoding. Because the encoder only ever sees a quarter of the patches, training runs more than 3 times faster; the method needs no human labels and still learns strong features, reaching 87.8% accuracy on ImageNet-1K with a ViT-Huge. In robotics it is commonly used to pretrain visual encoders: MVP pretrains an encoder with MAE and then freezes it for learning motor control, and VC-1 also uses an MAE objective.
ExampleMVP (2022) pretrains a ViT encoder with MAE on a large collection of real-world images, then freezes it and attaches a reinforcement-learning controller; on a set of arm-manipulation tasks it beats a supervised-pretrained encoder by up to 80 percentage points in success rate.
- Also called
- MAE, Masked Image Modeling
- Related
- Self-Supervised Learning · Vision Transformer · Pre-trained Visual Representation · MVP · VC-1 · Representation Learning
- Sources
- Masked Autoencoders Are Scalable Vision Learners (arXiv:2111.06377)
Masked Visual Pre-training for Motor Control (MVP, arXiv:2203.06173)