Embodied AI Glossary中文

R3M

Advanced

A general-purpose visual representation for robot manipulation, pretrained on first-person human videos.

R3M was released by Suraj Nair, Chelsea Finn, Abhinav Gupta, and colleagues at Stanford and Meta AI in March 2022, published at CoRL 2022. Robot demonstrations are scarce, which makes it hard to train a good visual encoder from scratch. R3M instead pretrains an image encoder on Ego4D, a large dataset of first-person human videos, with a combined objective: time-contrastive learning (frames close together in time should have similar representations), aligning video with its language description, and an L1 penalty that keeps the representation sparse and compact. The encoder is then frozen and used as a perception module, with only a policy trained on top for the downstream task. Across 12 simulated manipulation tasks, this raises success rates more than 20 points over training from scratch, and more than 10 points over CLIP or MoCo representations. R3M represents the 'pretrained visual representation' line of work and is often compared with VC-1, MVP, and VIP.

ExampleIn a cluttered real apartment, a Franka arm using a frozen R3M encoder as its visual module learned to manipulate objects with just 20 demonstrations per task.

Also called
R3M: A Universal Visual Representation for Robot Manipulation
Related
Pre-trained Visual Representation · VC-1 · MVP · VIP · Ego4D · Time-Contrastive Networks
Sources
R3M (arXiv 2203.12601)
As of
2022-03

See it in the full glossary →