Embodied AI Glossary中文

DINOv2

Common

Meta's 2023 self-supervised vision model that learns general-purpose image features with no labels at all.

DINOv2 is a vision foundation model Meta AI released in April 2023, the second generation of 2021's DINO, whose name stands for “self-distillation with no labels.” It uses no human annotation at all: a student network is trained to match a teacher network's output across different crops of the same image, a form of self-supervised learning, and the team also curated a 142-million-image dataset, LVD-142M, with an automated filtering pipeline. The largest model, ViT-g/14, has about 1.1 billion parameters, and was distilled down into smaller versions from 21 million to 300 million parameters. Its features need no fine-tuning — just attach a linear layer to do classification, segmentation, or depth estimation — and they preserve relatively fine spatial and geometric detail, exactly what robots need. OpenVLA fuses DINOv2's features with SigLIP's as its vision encoder, and DINO-WM trains a world model directly on DINOv2's patch features.

ExampleOpenVLA feeds the same camera image into both DINOv2 and SigLIP separately, fuses their features, and projects the result into the Llama 2 language model: DINOv2 supplies spatial and geometric detail, and SigLIP supplies semantics aligned with language.

Also called
self-DIstillation with NO labels v2, DINO Features
Related
Vision Foundation Model · Self-Supervised Learning · Vision Transformer · DINOv3 · Pre-trained Visual Representation · OpenVLA
Sources
DINOv2: Learning Robust Visual Features without Supervision (arXiv 2304.07193)
facebookresearch/dinov2 (GitHub)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)

See it in the full glossary →