Embodied AI Glossary中文

Figure Helix

HelixEssential

Figure AI's 2025 vision-language-action model for its humanoid robot, using a fast-slow two-system architecture to control the whole upper body.

Helix is a vision-language-action model (VLA — a model that looks at images, listens to instructions, and outputs actions directly) released on February 20, 2025, by the US humanoid-robot company Figure AI for use on its Figure humanoid robots. It uses a fast-slow, two-system design: System 2 is a 7-billion-parameter vision-language model that understands the scene and the language instruction at 7–9Hz and compresses the intent into a latent vector; System 1 is an 80-million-parameter Transformer that turns that latent vector, together with real-time perception, into actions at 200Hz. It controls the entire upper body at once — 35 degrees of freedom, including the wrists, every finger, the torso, and head orientation. Figure says training used only about 500 hours of teleoperation data, that a single set of weights covers many tasks without task-specific fine-tuning, and that the two systems each run on one of two low-power embedded GPUs onboard the robot. Later versions include Helix 02.

ExampleIn the launch demo, two Figure robots collaborated on putting away a bag of groceries neither had seen before, following spoken instructions and sorting items into the fridge and drawers.

Also called
Helix, Helix VLA
Related
Dual-System Architecture (System 1 / System 2) · Vision-Language-Action Model · Figure AI · Figure Helix 02 · Figure 02
Sources
Helix: A Vision-Language-Action Model for Generalist Humanoid Control (Figure AI)
As of
2025-02

See it in the full glossary →