Embodied AI Glossary中文

Pre-trained Visual Representation

预训练视觉表征PVRAdvanced

A vision encoder pretrained on large-scale images or video, reused as the 'eyes' for a robot policy.

A pre-trained visual representation is a vision encoder trained beforehand on large amounts of image or video data, which compresses a camera view into a feature vector for a robot policy or navigation agent to use, usually kept frozen or only lightly fine-tuned. It addresses the fact that robot data is scarce, making it inefficient to learn vision from scratch. In 2022, Parisi and colleagues found that representations pretrained only on generic vision data such as ImageNet let a control policy trained on top of them match or even beat one trained directly on ground-truth state, such as an object's exact position. Since then, PVRs built for embodied tasks have followed, including R3M (trained on Ego4D first-person human video), MVP (masked-autoencoder pretraining), and VC-1. The VC-1 paper compares these on CortexBench, a suite of 17 tasks, and finds no single PVR wins on all of them. VLAs today more often use vision foundation models like SigLIP or DINOv2 as the encoder directly instead.

ExampleVC-1 trains a ViT with masked autoencoding on more than 4,000 hours of first-person video plus ImageNet, then freezes it and attaches a small policy network, evaluated on CortexBench's locomotion, navigation, dexterous manipulation, and mobile manipulation tasks.

Also called
PVR, PVRs
Related
Vision Encoder · R3M · VC-1 · MVP · Vision Foundation Model · Backbone Freezing
Sources
The Unsurprising Effectiveness of Pre-Trained Vision Models for Control (arXiv 2203.03580)
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? (VC-1, arXiv 2303.18240)

See it in the full glossary →