Representation Alignment
表征对齐REPAAdvancedTraining a model so its intermediate features match those of an existing pretrained encoder.
REPA was proposed by Sihyun Yu, Saining Xie, and colleagues in October 2024 (an ICLR 2025 Oral), targeting the slow training of diffusion Transformers such as DiT and SiT. It adds an auxiliary loss: the network's intermediate hidden state while processing a noisy image, after passing through a small projection layer, is aligned with that same clean image's features from a pretrained visual encoder such as DINOv2. The authors argue that one bottleneck in diffusion-model training is having to learn good visual representations from scratch, and that borrowing an already-good representation saves that effort — SiT's training sped up by more than 17.5x. This “REPA-style” alignment has since been extended to video generation and robotics: Spatial Forcing, for example, aligns a VLA's intermediate visual tokens with features from the 3D foundation model VGGT, letting a VLA that has only ever seen 2D data implicitly pick up spatial awareness.
ExampleSpatial Forcing adds a cosine-similarity alignment loss against VGGT features on top of OpenVLA-OFT and π0, with no extra depth-map or point-cloud input, speeding up training by up to 3.8x and improving data efficiency.
- Also called
- REPA, REPA Regularization
- Related
- Representation Learning · Diffusion Transformer · DINOv2 · Auxiliary Loss / Auxiliary Task · VGGT · Pre-trained Visual Representation
- Sources
- Yu et al. 2024: Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
GitHub: sihyun-yu/REPA
Li et al. 2025: Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model - As of
- 2025-10