Embodied AI Glossary中文

ThinkAct

Advanced

A reasoning-based VLA that first has a multimodal large model work out a visual plan, then hands it to an action model to execute.

ThinkAct was released by NVIDIA and National Taiwan University (Chi-Pin Huang, Fu-En Yang, and others) in July 2025, published at NeurIPS 2025. Most VLAs map directly from an image and instruction to an action with no explicit reasoning step, which makes multi-step, long-horizon tasks difficult. ThinkAct uses a dual-system design: the 'thinking' part is a multimodal large model initialized from Qwen2.5-VL 7B, first given a supervised fine-tuning cold start, then trained with GRPO reinforcement learning to generate embodied reasoning plans, with rewards coming from visual signals — whether the endpoint of the planned trajectory matches the demonstration, and whether the whole trajectory is consistent with it. The reasoning output is compressed into a visual planning latent, which conditions a downstream action model that produces the specific actions. On manipulation benchmarks such as SimplerEnv and LIBERO, and on several embodied-reasoning benchmarks, it shows few-shot adaptation, long-horizon planning, and self-correction. A follow-up, Fast-ThinkAct, was published at CVPR 2026.

ExampleGiven 'put the carrot on the plate,' ThinkAct first reasons that it needs to grasp the carrot and then move it above the plate, and plans the gripper's motion trajectory, which the action model then executes; if the grasp fails partway through, it can replan.

Also called
Fast-ThinkAct (follow-up version), ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
Related
Embodied Reasoning · Dual-System Architecture (System 1 / System 2) · Chain-of-Thought · Group Relative Policy Optimization · Latent Reasoning · LIBERO Benchmark
Sources
ThinkAct (arXiv 2507.16815)
ThinkAct project page
Fast-ThinkAct (arXiv 2601.09708)
As of
2026-01

See it in the full glossary →