Embodied AI Glossary中文

Magma (Microsoft)

MagmaAdvanced

A multimodal agent foundation model from Microsoft that can both operate software interfaces and control a robot arm.

Magma is a multimodal foundation model released by Microsoft Research in February 2025, published at CVPR 2025, with the open-source version Magma-8B (language backbone Llama-3-8B). It aims to have one model handle agent tasks in both the digital and physical world: clicking through web pages and phone interfaces (UI navigation), and controlling a robot arm to pick and place objects. The key is two kinds of annotation: Set-of-Mark (SoM) numbers the actionable elements in an image, letting the model express where to act by “choosing which mark”; Trace-of-Mark (ToM) labels the future motion trajectory of marked points in a video, teaching the model to predict what happens next. With ToM, it can learn spatiotemporal planning from large amounts of instructional video that has no action labels at all. Training data includes UI-navigation data, Open X-Embodiment robot data, and web video.

ExampleThe same Magma-8B model can output “click the button marked 3” on a phone screenshot, and also output a robot arm's movement trajectory and gripper action on a robot camera feed.

Also called
Magma-8B, A Foundation Model for Multimodal AI Agents
Related
Visual Prompting (Set-of-Mark) · Vision-Language-Action Model · Multimodal Large Language Model · Open X-Embodiment · Llama · Microsoft Research
Sources
arXiv 2502.13130: Magma
GitHub: microsoft/Magma
As of
2025-02

See it in the full glossary →