Embodied AI Glossary中文

Contrastive Learning

对比学习Common

Learning features from unlabeled data by pulling similar samples' representations together and pushing dissimilar ones apart.

Contrastive learning is a family of self-supervised representation-learning methods: construct “positive pairs” (two pieces of data that should be similar, such as two random augmentations of the same image, or an image and its caption) and “negative” examples, then train an encoder to pull positive pairs' vectors closer together and push negatives further apart. A common loss for this is InfoNCE, from the 2018 CPC paper. Well-known examples include SimCLR, MoCo, and CLIP. Because it needs no human labels, it can learn general-purpose visual features from huge amounts of images, video, and image-text pairs. In embodied AI, many VLA models' vision encoders (such as CLIP and SigLIP) are trained with a contrastive objective; R3M instead pretrains robot visual representations on human videos using time-contrastive learning.

ExampleCLIP trains on 400 million image-text pairs scraped from the web, learning to tell which caption matches which image. R3M pretrains on Ego4D human videos with time-contrastive learning among other objectives, and once frozen and handed to a Franka arm, needs only 20 demonstrations to learn manipulation tasks in a real, cluttered apartment.

Also called
Contrastive Representation Learning
Related
Self-Supervised Learning · InfoNCE Loss · CLIP · SigLIP · Time-Contrastive Networks · R3M
Sources
Representation Learning with Contrastive Predictive Coding (arXiv 1807.03748)
Learning Transferable Visual Models From Natural Language Supervision (CLIP, arXiv 2103.00020)
R3M: A Universal Visual Representation for Robot Manipulation (arXiv 2203.12601)

See it in the full glossary →