Embodied AI Glossary中文

MOKA

MOKA(标记式视觉提示操作)Advanced

A training-free method that marks up an image for GPT-4V to pick keypoints, then converts those into robot-arm actions.

MOKA was released in March 2024 by Fangchen Liu, Kuan Fang, Pieter Abbeel, and Sergey Levine at UC Berkeley, published at RSS 2024. A vision-language model has plenty of common sense, but doesn't directly output robot actions. MOKA rewrites “how to manipulate this” as a picture-answering question: it overlays candidate points, a grid, and text labels onto the camera image as marks (visual prompting), and has GPT-4V hierarchically pick out a grasp point, the point where a tool contacts an object, a target point, and intermediate waypoints, which are then converted into robot-arm motion. It requires no robot-data collection for a new task, and can handle tabletop tasks such as tool use, deformable-object manipulation, and object rearrangement; successful experience gathered during execution can also serve as in-context examples, or be distilled into a policy network. It is a representative example of connecting a VLM to a robot through visual prompting.

ExampleGiven the instruction “use the brush to sweep the debris aside,” GPT-4V picks, on the marked image, the grasp point on the brush handle, the contact point where the bristles touch the table, and waypoints along the sweeping direction, and the robot executes them in that order.

Also called
Open-World Robotic Manipulation through Mark-Based Visual Prompting
Related
Visual Prompting (Set-of-Mark) · Affordance · Vision-Language Model · Zero-shot · Tool Use · PIVOT
Sources
MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting (arXiv 2403.03174)
MOKA project page
As of
2024-09

See it in the full glossary →