Embodied AI Glossary中文

Mobility VLA

Advanced

A hierarchical navigation system that has a long-context VLM watch a tour video, then follows an image-and-text instruction to find the destination.

This was released by Google DeepMind in July 2024. The task it targets is called MINT (Multimodal Instruction Navigation with demonstration Tours): someone first walks through the environment with a camera to record a tour video, and afterward the user can give instructions with text plus an image — for example, holding an object and asking “where does this go back.” The system has two layers: the high level uses Gemini 1.5 Pro, with a context length of up to 1 million tokens, to read through the entire tour video and the instruction and locate the frame where the target is; the low level uses COLMAP (a tool that recovers camera pose from images) to build a topological map from the video — a map where locations are nodes and passable connections are edges — and generates waypoint actions from it for the base to execute. In a real, occupied 836-square-meter office, end-to-end success rates for instructions requiring reasoning and for multimodal instructions were 86% and 90%, respectively.

ExampleThe user holds up a charger and asks “where should this go back,” and the robot first locates, within the tour video, the desk where the charger belongs, then drives there along the topological map.

Also called
MINT (Multimodal Instruction Navigation with demonstration Tours), Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs
Related
Vision-and-Language Navigation · Topological Map · Hierarchical Architecture · Google Gemini · Context Length · Structure from Motion
Sources
Mobility VLA (arXiv 2407.07775)
Mobility VLA (arXiv HTML full text)
As of
2024-07

See it in the full glossary →