Embodied AI Glossary中文

UBTech Thinker

优必选 ThinkerAdvanced

UBTech's in-house embodied model stack, made up of a foundation model, a world model, and an action model.

Thinker is the in-house embodied model system UBTech has built for its humanoid robots, reportedly organized into three layers: a foundation model called Thinker, a world model called Thinker-WM, and an action model called Thinker-VLA. Thinker itself is a vision-language model for embodied intelligence; UBTech open-sourced Thinker-4B (based on the Qwen3-VL architecture, 4B parameters, non-commercial license) along with a paper in January 2026. It targets problems that arise when a general vision-language model is applied to robots, such as viewpoint confusion and weak temporal understanding, and is trained on first-person video, visual grounding, spatial understanding, and chain-of-thought data, focused on task planning, visual grounding, and spatial understanding; the company states it leads on 7 embodied benchmarks. According to reports, Thinker-WM ranks first on the LIBERO benchmark, and Thinker-VLA raises inference efficiency by 176% in industrial settings. This model stack powers UBTech's industrial humanoid robots, including the Walker S series.

ExampleGiven a first-person image from a robot and the instruction 'put the part in the bin,' Thinker-4B can output a step-by-step task plan and also draw a box around the target part's location in the image.

Also called
Thinker-4B, Thinker-WM, Thinker-VLA
Related
UBTech Robotics · UBTech Walker S2 · Vision-Language Model · World Model · Vision-Language-Action Model · Qwen-VL
Sources
UBTECH-Robot/Thinker (GitHub)
Thinker: A vision-language foundation model for embodied intelligence (arXiv 2601.21199)
Embodied Intelligence 2026: Farewell to Narrative Hype, Practical Deployment Reigns Supreme (36Kr)
As of
2026-08

See it in the full glossary →