Embodied AI Glossary中文

InternVLA (Shanghai AI Laboratory)

上海AI实验室 InternVLA 系列Advanced

A family of embodied models from Shanghai AI Lab: M1 and A1 for manipulation, N1 for navigation.

InternVLA is a family of VLA (vision-language-action) models open-sourced by Shanghai AI Lab's embodied-AI team (GitHub organization InternRobotics). InternVLA-M1 (October 2025) is built on Qwen2.5-VL, first pretrained for spatial grounding on 2.3 million box, point, and trajectory annotations, then guided by spatial prompts to a diffusion-Transformer action expert that produces actions. InternVLA-A1 (January 2026, in 2B and 3B sizes) uses a Mixture-of-Transformers architecture that puts three experts — understanding, future-frame prediction, and action — into one model, pretrained on real-robot data, synthetic simulation data (such as InternData-A1), and human video, totaling 692 million frames. A1.5 (July 2026) switches to Qwen3.5-2B, and during training has a frozen Wan2.2 video model supervise “foresight tokens,” while inference generates no video at all. InternVLA-N1 is for navigation: a slow system marks waypoints on the image, and a fast system uses a diffusion policy to output trajectories at more than 30Hz.

ExampleWhen InternVLA-A1 sorts items off a moving conveyor belt, it has to predict where an object will be an instant later before grasping it; the current version of the paper reports it beats π0.5 by 26.7% on this kind of dynamic task.

Also called
InternVLA-M1, InternVLA-A1, InternVLA-A1.5, InternVLA-N1
Related
Vision-Language-Action Model · Dual-System Architecture (System 1 / System 2) · InternData-A1 · Vision-and-Language Navigation · Shanghai Artificial Intelligence Laboratory · World Action Model
Sources
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy (arXiv 2510.13778)
InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation (arXiv 2601.02456)
InternVLA-N1 项目主页 (Chinese)
As of
2026-07

See it in the full glossary →