Embodied AI Glossary中文

MemoryVLA

Advanced

Adds working memory and a memory bank to a VLA, so the robot remembers what it has already seen and done.

MemoryVLA is work from Gao Huang's group at Tsinghua University, in collaboration with Dexmal (原力灵机), Megvii, and others, released in August 2025 and published at ICLR 2026. Most VLA models decide on an action by looking only at the current frame, but for many manipulation tasks a single current frame isn't enough — the footage right before and right after pressing a button can look nearly identical, and without any memory of history there's no way to tell whether it's already been pressed. MemoryVLA borrows the cognitive-science concepts of working memory and episodic memory: a 7B vision-language model encodes observations into perceptual tokens and cognitive tokens, serving as working memory; a separate memory bank stores past low-level detail and high-level semantics, retrieved as needed and fused with current information; a diffusion action expert then outputs a segment of action. The paper reports a 71.9% success rate on SimplerEnv-Bridge, 96.5% on LIBERO, and 84.0% on real-robot tasks, with especially large gains on long-horizon tasks.

ExampleFor a task like “press three buttons in sequence,” the footage barely changes after each press, so only the memory bank remembering which ones are already pressed can prevent the robot from pressing the same one twice or skipping one.

Also called
Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
Related
Memory-Augmented VLA · Embodied Memory · Vision-Language-Action Model · Long-horizon Task · Diffusion Action Head · Dexmal
Sources
arXiv 2508.19236: MemoryVLA
MemoryVLA 项目主页 (Chinese)
As of
2026-01

See it in the full glossary →