MolmoAct
AdvancedAi2's open-source “action reasoning model,” a VLA that reasons about depth and sketches a trajectory before outputting an action.
MolmoAct was released by Ai2 in August 2025, built on the company's own Molmo vision-language model, with 7 billion parameters; the model, code, data, and evaluation scripts are all open-source. It splits VLA decision-making into three steps: first it outputs perceptual tokens carrying depth information to understand 3D space, then it draws a waypoint trajectory for the end effector directly on the image, and finally it decodes that trajectory into specific arm and gripper actions. Because the planned trajectory is overlaid directly on the image, a person can see what it intends to do before execution, and can also draw a path on a phone or tablet to guide it. At release, it reported a 72.1% success rate on SimplerEnv out-of-distribution tasks and 86.6% on LIBERO. MolmoAct2, from May 2026, switched to the embodied-reasoning backbone Molmo2-ER and attached a flow-matching action expert, cutting single-action-inference time from about 6.7 seconds down to 180 milliseconds, and released a companion dataset of more than 720 hours of bimanual YAM data.
ExampleGiven the instruction “put the pillow on the sofa,” MolmoAct first draws a trajectory from pillow to sofa on the camera view; after the user confirms it or redraws it on a tablet, it converts that into robot-arm actions to execute.
- Also called
- MolmoAct2, Action Reasoning Model (ARM), Action Reasoning Models that can Reason in Space
- Related
- Vision-Language-Action Model · Molmo (Ai2) · Embodied Reasoning · Action Chain-of-Thought · Flow Matching · MolmoAct2-BimanualYAM
- Sources
- MolmoAct (Ai2 blog)
MolmoAct: Action Reasoning Models that can Reason in Space (arXiv 2508.07917)
MolmoAct 2 (Ai2 blog) - As of
- 2026-05