Memory-Augmented VLA
记忆增强 VLAAdvancedA VLA with an added history-memory module, so it can act based on what happened earlier, not just the current frame.
Memory-augmented VLA is an umbrella term for a research direction, not one specific model. Most mainstream VLAs (vision-language-action models) only look at the current frame or the last few frames to produce an action, implicitly assuming the current view holds all the information decisions need. But many manipulation tasks depend on history: a button looks the same whether or not it's already been pressed, and an object hidden behind something else isn't visible in the frame at all. Stuffing a long history of frames straight into the model would blow up the token count and inference time. This line of work instead maintains a separate memory: past observations are compressed into features and stored in a memory bank, and at decision time relevant entries are retrieved and fused with the current features before going to the action head. Representative work includes MemoryVLA, proposed in 2025 by Tsinghua and Yuanli Robotics, among others; the MIKASA-Robo benchmark specifically tests this kind of memory ability across 32 tabletop manipulation tasks.
ExampleIn MemoryVLA's real-robot task 'press three buttons in a specified color order,' a pressed and an unpressed button look identical, so progress can't be judged from the current frame alone; MemoryVLA completes the task using history stored in its memory bank, and the paper reports it beats the strongest baseline by 26 points on long-horizon tasks.
- Related
- Memory-Augmented VLA · Vision-Language-Action Model · Embodied Memory · Long-horizon Task · Partially Observable Markov Decision Process · History Encoder
- Sources
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation (arXiv:2508.19236)
Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning (MIKASA, arXiv:2502.10550) - As of
- 2026-03