VIMA-Bench
AdvancedA tabletop manipulation benchmark driven by interleaved text-and-image prompts, with four levels of generalization testing.
VIMA-Bench was introduced alongside the VIMA model, first-authored by Yunfan Jiang with collaborators including Fei-Fei Li and Anima Anandkumar, published at ICML 2023. Built on the Ravens simulator (a PyBullet-based tabletop manipulation environment), it extends to 17 meta-tasks whose instructions are written as multimodal prompts interleaving text and images, spanning categories such as simple object manipulation, visual goal reaching, novel-concept understanding, one-shot video imitation, visual-constraint satisfaction, and visual reasoning, and it provides 650,000 successful demonstration trajectories. Evaluation has four levels: L1 only randomizes object placement; L2 recombines already-seen objects and textures in new ways; L3 introduces new objects and textures; L4 is entirely new tasks — 4 of the 17 tasks are held out specifically to test zero-shot generalization. The action space is one grasp pose plus one placement pose.
ExampleThe VIMA paper reports that under the hardest zero-shot generalization setting, with the same amount of training data VIMA's task success rate was up to 2.9 times that of other approaches, and still 2.7 times higher with 10 times less training data.
- Also called
- VIMABench
- Related
- VIMA · Benchmark · Compositional Generalization · Zero-shot · Tabletop Manipulation · Transporter Networks
- Sources
- VIMA: General Robot Manipulation with Multimodal Prompts (arXiv 2210.03094)
vimalabs/VIMABench GitHub 仓库 (Chinese) - As of
- 2023-05