Latent Reasoning
潜在推理AdvancedHaving a model carry out multi-step reasoning inside its internal hidden states, instead of writing out a chain of thought in words.
Chain-of-thought has a large model write out intermediate reasoning steps in words before giving an answer; it works well, but generating a token at every step is slow and constrained by language. Latent reasoning instead keeps the intermediate steps inside the model's continuous hidden state: Coconut, proposed by Meta and others in December 2024, treats the last layer's hidden state as one step of 'continuous thought,' feeding it directly back in as the next step's input embedding instead of decoding it into words, and outperforms word-based chain-of-thought on logic problems that need a lot of search, while generating fewer tokens too. A July 2025 survey defines this direction as reasoning done entirely in continuous hidden states, with no per-token supervision. Embodied AI needs both reasoning and low latency for a VLA, so it borrows the same idea: for example, ThinkAct compresses the reasoning plan a multimodal large model produces into a single 'visual plan latent,' which then guides a downstream action model. The cost is that the intermediate process becomes unreadable, making it hard to inspect or debug.
ExampleSolving a logic problem that needs several steps of deduction, Coconut doesn't write out text like 'because A, therefore B'; instead it feeds a hidden vector back into itself as input over several steps, only outputting the answer at the end.
- Also called
- Latent CoT, Continuous Thought
- Related
- Chain-of-Thought · Embodied Chain-of-Thought · Reasoning · Large Language Model · ThinkAct · Inference Latency
- Sources
- Training Large Language Models to Reason in a Continuous Latent Space (Coconut, arXiv:2412.06769)
A Survey on Latent Reasoning (arXiv:2507.06203)
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning (arXiv:2507.16815) - As of
- 2025-07