Embodied Memory
具身记忆AdvancedA robot storing what it has seen and done so it can use that later for decisions and answering questions.
Embodied memory refers to the mechanisms an embodied agent uses, over long periods of operation, to store and recall past experience: where it has been, where things are kept, how far a task has progressed, where it failed last time. This is necessary because a robot only ever sees part of the world at once, since it is partially observable, and many tasks run anywhere from minutes to days, so a policy that only looks at the current frame will forget sub-tasks it already finished, or keep searching for the same object over and over. There are roughly three common approaches. One keeps historical features inside the model itself — MemoryVLA (2025), for example, draws on human working memory and episodic memory to add a memory bank to a VLA model. Another maintains an external structured memory, such as a 3D scene graph or semantic map — KARMA uses a long-term scene graph plus short-term state records to support household task planning. The third is retrieval augmentation, as in ReMEmbR, which stores long-horizon navigation footage in a database indexed by time and location, used to answer questions about where and when something happened.
ExampleA patrol robot asked, “where did you last see a red cart,” has to retrieve the time and location from hours of historical footage before it can answer.
- Related
- Memory-Augmented VLA · Memory-Augmented VLA · Long-horizon Task · Partially Observable Markov Decision Process · 3D Scene Graph · Embodied Question Answering
- Sources
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
ReMEmbR: Building and Reasoning Over Long-Horizon Spatio-Temporal Memory for Robot Navigation
KARMA: Augmenting Embodied AI Agents with Long-and-short Term Memory Systems - As of
- 2025-08