VIMA
AdvancedA Transformer robot agent that uses interleaved text-and-image 'multimodal prompts' to describe manipulation tasks in one unified format.
VIMA was released in October 2022 by researchers at Stanford, NVIDIA, Caltech, and other institutions (Yunfan Jiang, Linxi Fan, Yuke Zhu, Fei-Fei Li, and others), published at ICML 2023. Robot tasks can be specified in many different ways — showing a demonstration to imitate, describing it in language, or giving a goal image. VIMA unifies all of these into a single 'text interleaved with images' multimodal prompt, such as 'put [image of an object] into [image of a container],' with a Transformer reading the prompt and autoregressively outputting actions. The authors also built the VIMA-Bench simulation benchmark: 17 task templates that can procedurally generate thousands of tabletop tasks, more than 600,000 expert trajectories, and a four-level evaluation of increasingly hard generalization. The paper reports up to 2.9x higher success than other designs in the hardest zero-shot setting.
ExampleA VIMA prompt is a sentence with two small images embedded in it: 'put [image of a red block] onto [image of a green plate]'; VIMA finds the matching objects on a simulated tabletop and completes the pick-and-place, and the same model also handles a prompt that first shows a demonstration image and then says 'do it like this.'
- Also called
- VIMA: General Robot Manipulation with Multimodal Prompts
- Related
- VIMA-Bench · Language-conditioned Policy · Compositional Generalization · Tabletop Manipulation · Imitation Learning · Transformer
- Sources
- VIMA (arXiv 2210.03094)
VIMA project page - As of
- 2023-05