Embodied AI Glossary中文

Representation Collapse

表征坍缩Advanced

When a self-supervised encoder outputs the same or nearly the same vector for every input, so the representation carries no information.

Representation collapse is a classic failure mode in self-supervised learning and joint-embedding predictive models. If the training objective only requires the representations of two views of the same sample (or of the current moment and the future) to be close to each other, the easiest solution for the encoder is to output the same constant vector for every input: the loss is tiny, but the representation carries no information at all. A milder form is called dimensional collapse, where the vector only makes use of a handful of dimensions. Anti-collapse methods fall roughly into three categories: contrastive learning uses negative samples to push different images' representations apart; VICReg (Bardes, Ponce, LeCun, 2021) explicitly constrains the variance of each dimension and the covariance between dimensions; and BYOL and the JEPA family use a stop-gradient plus an exponential-moving-average (EMA) target encoder. In embodied AI, training video representations or world models in a JEPA-style latent space, such as Meta's V-JEPA 2, has to deal with the same problem.

ExampleWhen training V-JEPA 2, the target representation for a masked video segment is computed by an EMA copy of the encoder's weights and has stop-gradient applied to it, which the paper explains is specifically meant to prevent representation collapse.

Also called
Dimensional Collapse
Related
Self-Supervised Learning · Joint-Embedding Predictive Architecture · Contrastive Learning · Stop-Gradient · Exponential Moving Average · V-JEPA 2
Sources
VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning (arXiv 2105.04906)
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (arXiv 2506.09985)
As of
2025-06

See it in the full glossary →