12Landmark Models & Projects
Famous models in historical order: from SayCan and RT-2 to the π series, GR00T, and world models. · 332 terms
- 12.1Foundation models embodied AI borrows11
- 12.2LLMs as the brain27
- 12.3End-to-end learning and the RT series16
- 12.4Classic imitation-learning policies21
- 12.5Open-source generalist policies and VLA32
- 12.6The π series and reinforcement learning28
- 12.7Global tech giants and star startups35
- 12.8Embodied models in China45
- 12.9World models and learning from video43
- 12.10Legged, dexterous-hand, and agile skills16
- 12.11Humanoid whole-body control and teleop37
- 12.12Navigation and autonomous driving21
12.1Foundation models embodied AI borrows
An embodied model’s ‘eyes’ and ‘brain’ are often borrowed: meet these foundation models and vision encoders first.
- OpenAI GPT SeriesGPT 系列(GPT-4o / GPT-5)OpenAI's large language model family; from GPT-4o onward it handles images and audio directly, often used by robots for high-level planning.
- A 2022 DeepMind vision-language model that learns a new task from just a few image-text examples, an early landmark VLM.
- LLaVALLaVA(视觉指令微调架构)A 2023 open-source multimodal model that connects a vision encoder to an LLM through a projection layer, then fine-tunes on instructions.
- A roughly 55-billion-parameter multilingual vision-language model from Google, one of the two backbones behind RT-2.
- Google GeminiGemini 系列(谷歌多模态大模型)Google DeepMind's family of natively multimodal large models, also the foundation for robotics models such as Gemini Robotics.
- InternVL书生 InternVLShanghai AI Lab's open-source vision-language model series, starting from a 6-billion-parameter vision encoder and scaling up from there.
- Molmo (Ai2)MolmoAn Ai2 vision-language model open-sourced with both weights and training data that can answer questions by “pointing” on the image.
- MVPMVP(掩码视觉预训练)Uses a masked autoencoder to pretrain a vision encoder on huge amounts of natural images, then freezes it for robot use.
- A general-purpose visual representation for robot manipulation, pretrained on first-person human videos.
- VIPVIP(价值隐式预训练)A self-supervised pretraining method on human videos that produces both a visual representation and a dense reward at once.
- Meta's visual encoder for embodied tasks, pretrained with masked autoencoding on more than 4,000 hours of first-person video.
12.2LLMs as the brain
The earliest use of large models was as the brain: breaking down tasks, writing code and rewards, understanding space, then calling existing skills.
- Combines what a language model says is useful with what a value function says is achievable to pick each action.
- Socratic ModelsSocratic Models(苏格拉底模型)A framework that chains several off-the-shelf large models together zero-shot using natural language as the common interface for multimodal tasks.
- Feeds environment feedback back into a large language model as text, letting the robot adjust its plan as it goes.
- Code as Policies代码即策略Has a large language model write Python code directly, calling perception and control APIs to command a robot.
- A method that prompts a large language model with Python-like code to generate robot task plans the robot can actually execute.
- Google's 2023 embodied multimodal large language model, which feeds images and robot-state estimates directly into PaLM.
- A 2023 embodied multimodal model from HKU and Shanghai AI Lab that uses chain of thought to generate step-by-step plans.
- A GPT-4-driven, lifelong-learning agent in Minecraft that sets its own goals and builds up a library of skills by writing code.
- KnowNoKnowNo(会求助的机器人)Has an LLM planner ask a person when uncertain, using conformal prediction to give a statistical success guarantee.
- A method that uses a 3D scene graph to let a large language model do long-horizon task planning across a large, multi-floor space.
- A large language model that takes 3D scene features directly as input to answer spatial questions and break down tasks.
- LEO (BIGAI)LEO(3D 具身通才智能体)A 3D embodied generalist model from BIGAI that understands 3D scenes and can answer questions, navigate, and manipulate objects.
- Language to RewardsLanguage to Rewards(L2R)Has a large language model translate plain instructions into a reward function, handed to a real-time optimizer to generate robot motion.
- Has GPT-4 write reward-function code and repeatedly refine it, automating the reward design that reinforcement learning depends on.
- Uses a large language model to automatically write reward functions and domain-randomization ranges, transferring a simulation-trained policy to a real robot.
- A zero-shot method where a large model writes code to build a 3D value map, which a planner turns into a trajectory.
- A 2023 MIT method that distills 2D features like CLIP into a 3D field, letting a robot grasp objects from language instructions with few demonstrations.
- A 2024 NYU pick-and-place system built by combining off-the-shelf vision-language models with navigation and grasping modules.
- A method that teaches a vision-language model to estimate distance and size using a massive, automatically generated set of 3D spatial question-answer pairs.
- PIVOTPIVOT(迭代视觉提示)A method that controls robots zero-shot by drawing candidate actions on an image and having a VLM repeatedly pick and narrow them down.
- MOKAMOKA(标记式视觉提示操作)A training-free method that marks up an image for GPT-4V to pick keypoints, then converts those into robot-arm actions.
- A training-free framework where GPT-4V finds relevant object parts and writes spatial constraints to follow open-ended manipulation instructions.
- A vision-language model that points to where to place or grasp something on an image, following a language instruction.
- Has a large model write task constraints between keypoints as code, then solves for robot motion with optimization.
- A zero-shot manipulation framework that turns a vision-language model's reasoning into point-and-direction constraints defined in each object's own coordinate frame.
- SpatialLMSpatialLM(群核空间大模型)Manycore's open-source 3D large model that reads an indoor point cloud and outputs a structured layout of walls, doors, windows, and furniture.
- A 3B embodied-reasoning model from Tianjin University that uses “pointing” as an intermediate representation, trained with reinforcement fine-tuning.
12.3End-to-end learning and the RT series
The other path is learning control end to end: from real-robot grasping to generalist models, up to RT-2, which coined VLA.
- End-to-End Training of Deep Visuomotor Policies端到端视觉运动策略(引导策略搜索)A landmark 2015 Berkeley paper that used a convolutional network to output robot-arm joint torques directly from camera images.
- Google Arm FarmGoogle 机械臂农场(大规模抓取自监督)Google's 2016 project where more than a dozen robot arms attempted grasps over 800,000 times to learn grasping from the data.
- Dex-Net 2.0Dex-Net(GQ-CNN 抓取网络)Berkeley's grasp-scoring network, trained on simulated data, that judges from a depth image which grasp will hold best.
- Google's large-scale reinforcement learning grasping system, trained on more than 580,000 real-robot grasp attempts to learn a visual Q-function.
- A Google multi-task reinforcement-learning system that used 7 real robots to gather 9,600 hours of data while learning 12 tasks at once.
- Decision Transformer决策 TransformerTreats reinforcement learning as sequence modeling: given a target return, a GPT-style model predicts actions one step at a time.
- A 2022 DeepMind generalist agent whose single set of weights can play games, caption images, chat, and control a robot arm.
- A Transformer robot agent that uses interleaved text-and-image 'multimodal prompts' to describe manipulation tasks in one unified format.
- Google's 2022 Transformer policy for real robots, trained on 130,000-plus real-robot demonstrations across 700-plus tasks.
- Google DeepMind's 2023 model that outputs robot actions as text tokens — the paper that coined the term VLA.
- RT-1 and RT-2 retrained on Open X-Embodiment, a large dataset pooling robot data across 22 different robot embodiments.
- A policy that replaces language instructions with a trajectory sketch drawn on the image, showing the robot how the motion should go.
- DeepMind's 2024 system that uses a large model to automatically assign tasks to a fleet of robots and collect real data.
- A hierarchical VLA from Google that has the robot first state a 'language motion' like 'move arm forward' before outputting the action.
- DeepMind's cross-embodiment manipulation agent, built on Gato, that generates its own data to keep improving.
- An early piece of work that fine-tuned the open-source OpenFlamingo vision-language model directly into a robot manipulation policy.
12.4Classic imitation-learning policies
Smaller imitation-learning models were evolving in parallel: from PerAct to Diffusion Policy and ACT, mastering fine bimanual work.
- A highly sample-efficient manipulation network that treats pick-and-place as 'moving one region of the image to another location.'
- A language-conditioned manipulation policy that combines CLIP's semantic understanding with Transporter Networks' pixel-level spatial precision.
- Neural Descriptor Fields神经描述子场An MIT 2021 object representation that transfers a manipulation skill to new objects of the same category, in new poses, from just a few demonstrations.
- A multi-task manipulation policy that voxelizes the scene and uses a Perceiver Transformer to predict the next key pose.
- Motion Policy Networks运动策略网络A neural motion planner that generates collision-free robot-arm motion directly from a single depth camera's point cloud.
- Behavior TransformerBeT / VQ-BeTA Transformer-based imitation-learning method that learns several different valid behaviors at once from multimodal demonstration data.
- DiffuserDiffuser(扩散规划器)A planning method that generates an entire trajectory as one object to be denoised, one of the first uses of diffusion for decision-making.
- Decision DiffuserDecision Diffuser(决策扩散器)Generates future state trajectories with a return-conditioned diffusion model, then infers actions from them, skipping dynamic programming.
- Diffusion Policy扩散策略An imitation-learning policy that generates robot action sequences by gradually denoising from random noise, using a diffusion model.
- A 2023 Stanford imitation-learning policy that predicts a whole chunk of actions at once, letting low-cost bimanual robots do fine manipulation.
- RoboAgentRoboAgent(MT-ACT)A multi-skill kitchen manipulation agent from CMU and Meta, trained on just 7,500 demonstrations.
- Dobb-EDobb·EAn open-source household-robot framework from NYU and Meta, released in 2023, that teaches a robot a new task from just 5 minutes of demonstration.
- Stanford's 2024 project that puts the ALOHA bimanual robot on a mobile base for low-cost whole-body teleoperation and imitation learning.
- DeepMind's 2024 work: massive teleoperated data plus a diffusion policy teach a bimanual robot to tie shoelaces.
- An imitation-learning policy that combines diffusion policies with 3D scene representations to generate robot-arm end-effector trajectories.
- 3D Diffusion Policy3D 扩散策略A diffusion policy that takes sparse point clouds as input and learns manipulation from very few demonstrations.
- An improved 3D Diffusion Policy that lets a humanoid trained on data from a single scene generalize to new scenes.
- Consistency PolicyConsistency Policy(一致性策略)Distills a diffusion policy into a fast visuomotor policy that generates an action in a single step.
- NVIDIA's multi-view 3D manipulation policy that learns millimeter-precision insertion from about 10 demonstrations per task.
- Policies that perform single tasks such as opening a cabinet or a drawer in unfamiliar homes with no fine-tuning at all.
- A system that places low-cost tactile-sensor readings and a visual point cloud in the same 3D space to learn fine manipulation.
12.5Open-source generalist policies and VLA
RT and imitation learning converge into open-source VLA: after Octo and OpenVLA, improvements bloomed in every direction.
- A 2024 open-source generalist robot policy pretrained on 800,000 cross-embodiment trajectories, quick to fine-tune to new robots.
- A 7-billion-parameter vision-language-action model that Stanford, UC Berkeley, and collaborators open-sourced in 2024, code and weights included.
- Stanford's fine-tuning recipe for VLAs that makes OpenVLA generate actions 26 times faster, with a higher success rate.
- A single set of weights that controls single arms, bimanual arms, quadrupeds, ground vehicles, and drones with one cross-embodiment Transformer.
- Heterogeneous Pre-trained Transformers异构预训练 TransformerA 2024 method from Kaiming He's group at MIT that pretrains one shared backbone jointly across many different robots' data.
- Tsinghua's 2024 diffusion foundation model for bimanual manipulation, with about 1.2 billion parameters and a unified action space.
- A VLA foundation model from Tsinghua, trained on tens of thousands of hours of UMI data, that deploys zero-shot on new robot arms.
- An efficient VLA that replaces the Transformer language backbone with the Mamba state-space model.
- A fast VLA that pairs a small multimodal model with a diffusion policy head and skips robot-data pretraining altogether.
- An open-source VLA that attaches a diffusion Transformer action module behind a frozen vision-language backbone.
- A method that draws the recent motion trace of key points onto the image as a prompt, boosting a VLA's sense of space and time.
- A systematic study and open-source framework comparing which VLM backbone, architecture, and training data work best for building a VLA.
- UniActUniAct(通用动作空间)A method for training an embodied foundation model on a shared, discrete set of 'universal actions' common across different robots.
- HAMSTERHAMSTER(分层动作模型)A 2025 hierarchical VLA from NVIDIA and others where a high level sketches a 2D path and a low-level 3D policy follows it.
- A VLA from Midea and others that attaches a roughly billion-parameter diffusion action expert to a vision-language model, adapting to many robots.
- Magma (Microsoft)MagmaA multimodal agent foundation model from Microsoft that can both operate software interfaces and control a robot arm.
- A unified VLA model that can both answer questions about images and directly control a robot.
- A hierarchical dexterous-hand grasping framework from Peking University and PsiBot: a large model plans, a diffusion policy acts.
- A VLA that predicts actions both by diffusion and autoregressively, inside the same large language model.
- An open-source, roughly 330-million-parameter generalist VLA policy that denoises a whole action sequence directly with a single Transformer.
- A 7-billion-parameter VLA that first generates a future subgoal image as its “thought,” then outputs an action to reach it.
- Hugging Face's open-source, roughly 450-million-parameter small VLA that trains and runs on ordinary consumer hardware, even a laptop.
- An action tokenizer that compresses a chunk of actions into a fixed number of tokens using B-spline control points.
- A 3D-manipulation VLA that projects point clouds into 2D images and reads off actions as heatmaps on them.
- A VLA model that first predicts future dynamic regions, depth, and semantic features, then generates actions conditioned on them.
- A reasoning-based VLA that first has a multimodal large model work out a visual plan, then hands it to an action model to execute.
- Ai2's open-source “action reasoning model,” a VLA that reasons about depth and sketches a trajectory before outputting an action.
- Adds working memory and a memory bank to a VLA, so the robot remembers what it has already seen and done.
- Discrete Diffusion VLADiscrete Diffusion VLA(离散扩散 VLA)Decodes VLA actions with discrete diffusion: action tokens are generated in parallel, confident ones fixed first, the rest filled in over several rounds.
- A VLA that reaches top-tier performance without any robot pretraining, using a 0.5B small model plus a lightweight policy module.
- A 0.9B cross-embodiment VLA that unifies many robots by giving each one its own learnable 'soft prompt.'
- NVIDIA's minimalist VLA that makes no architecture changes at all, having the VLM write out actions directly as plain digit text.
12.6The π series and reinforcement learning
Physical Intelligence’s π series set the VLA benchmark; then how reinforcement learning makes policies stronger from experience.
- Physical Intelligence's 2024 VLA that generates continuous robot actions with flow matching; code and weights are open-sourced.
- An autoregressive VLA that tokenizes π0's actions into discrete tokens with the FAST tokenizer and predicts them one by one.
- PI's 2025 VLA that handles long-horizon tasks like tidying a kitchen in real homes it has never seen.
- π0.6 after reinforcement learning with RECAP, able to keep improving from demonstrations, corrections, and its own experience.
- PI's 2026 generalist robot model that can be steered with rich prompts and shows early signs of compositional generalization.
- A 2025 Physical Intelligence hierarchical system where a high-level VLM breaks down complex instructions and low-level π0 executes them.
- Physical Intelligence's method for letting a VLA keep improving from its own experience plus human corrections.
- A large-scale offline reinforcement learning method that uses a Transformer to estimate Q-values one action dimension at a time.
- An open-source real-world reinforcement learning toolkit from Berkeley and others that can train a policy on a real robot arm in tens of minutes.
- Berkeley's real-robot RL system where a human can take over anytime, learning fine manipulation in 1 to 2.5 hours.
- Diffusion Policy Policy OptimizationDPPO(扩散策略策略优化)Fine-tunes a diffusion policy directly with PPO policy-gradient reinforcement learning by treating each denoising step as a decision.
- A method that re-ranks a generalist policy's candidate actions at deployment time using a value function learned with offline reinforcement learning.
- Generative Value Learning (GVL)GVL(生成式价值学习)Has a vision-language model estimate task progress from shuffled video frames, turning it into a general-purpose value function.
- A method that uses preference alignment on successful and failed trajectories to improve a VLA's generalization to new tasks.
- A method that fine-tunes a VLA with reinforcement learning using a consistency policy, offline first, then online on the real robot.
- RIPT-VLARIPT-VLA(交互式后训练)A method that post-trains a pretrained VLA with reinforcement learning using only a binary success/failure reward.
- An early framework that further improves an autoregressive VLA, such as OpenVLA, using online reinforcement learning.
- A method that injects learnable noise into a flow-matching policy so it can be fine-tuned with online reinforcement learning.
- Diffusion Steering via Reinforcement LearningDSRL(扩散策略噪声空间强化学习)Leaves a diffusion policy's weights untouched and uses reinforcement learning only to pick its input noise, steering it toward better actions.
- A method that samples several candidate action sets at deployment time and uses a VLM-based verifier to pick the best one, improving VLA robustness.
- A method that scales long-horizon task data by having a human take over right before failure, first recovering and then correcting.
- An open-source framework that runs large-scale online reinforcement learning on a VLA using only a success/failure reward.
- A method that runs reinforcement learning inside a learned world model to make a VLA more robust in just a few hundred steps.
- A diffusion-policy-based real-world reinforcement learning framework that reached 1,000-for-1,000 success across eight manipulation tasks.
- An open-source framework for online reinforcement-learning fine-tuning of flow-matching VLAs like π0 and π0.5.
- WMPOWMPO(基于世界模型的策略优化)A method that lets a VLA do reinforcement learning on trajectories 'imagined' by a video world model, with no real-robot interaction.
- A VLA post-training method where a strong teacher corrects the student token by token on trajectories the student generated itself.
- A general-purpose robot reward model trained on both task-progress labels and pairwise trajectory comparisons.
12.7Global tech giants and star startups
From technical approaches to companies: the foundation models of NVIDIA, Figure, Google, and other firms outside China.
- NVIDIA Isaac GR00T N1GR00T N1 系列NVIDIA's open humanoid-robot foundation model family, built on a fast-slow architecture and continually updated since its 2025 launch.
- NVIDIA Isaac GR00T N2GR00T N2NVIDIA's next-generation robot foundation model, previewed for 2026, moving from a VLA design to a world-action-model architecture.
- NVIDIA CosmosCosmosNVIDIA's open world-foundation-model platform, using video generation to produce training data and rehearsal environments for robots and self-driving cars.
- NVIDIA Cosmos PredictCosmos PredictThe video world model branch of NVIDIA's Cosmos family, generating what comes next from text, an image, or video.
- NVIDIA Cosmos TransferCosmos TransferNVIDIA's controllable video-generation model that produces photorealistic video conditioned on structure maps like depth, segmentation, and edges.
- NVIDIA Cosmos ReasonCosmos ReasonNVIDIA's reasoning vision-language model for physical AI, which watches a video, reasons through it in writing, then answers.
- Fine-tunes a video generation model directly into a robot policy that predicts actions, future frames, and success confidence together.
- NVIDIA's third-generation omnimodal world model: a single model that understands video, generates it, and also outputs actions.
- DreamGenDreamGen(GR00T Dreams)An NVIDIA 2025 method that uses a video world model to generate robot videos and infer actions from them to synthesize training data.
- A generalist robot world model NVIDIA released in 2026, pretrained on about 44,000 hours of human video.
- NVIDIA's 2026 world action model, built on a 14B video diffusion model that predicts frames and actions together.
- An NVIDIA 2026 project that pretrains a dexterous-hand VLA on 20,000 hours of first-person human video.
- Figure HelixHelixFigure AI's 2025 vision-language-action model for its humanoid robot, using a fast-slow two-system architecture to control the whole upper body.
- Figure Helix 02Helix 02Figure's 2026 full-body humanoid control model, mapping raw pixels directly to whole-body control for long-horizon tasks.
- A humanoid robot model Figure AI released in September 2026 that does household chores zero-shot in 30 unfamiliar homes.
- 1X World Model1X 世界模型Humanoid company 1X's video world model, first used to evaluate policies, later used directly to control NEO.
- Redwood1X Redwood1X's vision-language-action model for its home humanoid NEO, running entirely on the robot's onboard GPU.
- Google DeepMind's 2025 VLA built on Gemini 2.0, able to directly control robots for dexterous manipulation.
- Google DeepMind's embodied-reasoning model, which understands spatial layouts, makes plans, and judges whether a robot task is done.
- Google DeepMind's lightweight Gemini Robotics VLA that runs directly on the robot itself, with no internet connection required.
- Google DeepMind's VLA that thinks in words before acting, with skills that transfer across different kinds of robots.
- Google DeepMind's 2026 robotics model generation, with the new ability to control a humanoid's whole body.
- Veo World SimulatorVeo 世界模拟器A Google DeepMind system that uses the Veo video model to 'imagine' a robot's execution in order to evaluate policies.
- Rho-alpha微软 Rho-alphaMicrosoft's first robotics model, built from its Phi vision-language models and adding touch sensing.
- RFM-1Covariant RFM-1Covariant's roughly 8-billion-parameter robot foundation model, trained on warehouse picking data.
- Dyna Robotics' first-generation robot foundation model, released in 2025, best known for folding napkins autonomously for hours at a stretch.
- DYNA-2Dyna DYNA-2Dyna Robotics' 2026 world action model, pretrained on a million hours of first-person human video.
- Skild AI's general-purpose robot brain, with a single model able to control quadrupeds, humanoids, arms, and other embodiments.
- Skild AI's robot foundation model that can perform a task it never trained on after watching a single demonstration video.
- Field Foundation Models (Field AI)Field AI FFM(场景基础模型)Field AI's robot foundation model, built to operate autonomously in unmapped construction sites and industrial facilities.
- Generalist AI's embodied foundation model, trained on more than 270,000 hours of real-world interaction data from homes and workplaces.
- Generalist AI's second-generation embodied foundation model, aiming for reliable, fast “mastery” of simple tasks rather than just getting them done.
- Sunday Robotics ACT-1Sunday ACT-1Sunday Robotics' 2025 household robot model, trained with not a single teleoperated robot demonstration in its data.
- GENE-26.5Genesis AI GENE-26.5Genesis AI's first robotics foundation model system, released in May 2026, focused on dexterous manipulation.
- Isaac 0.5 (Perceptron)Perceptron Isaac 0.5A 36-billion-parameter open-weight robot foundation model from Perceptron, released in 2026, unifying seeing, thinking, and acting.
12.8Embodied models in China
Now China: foundation models from big tech, research institutes, and startups, mostly split between VLA and world-action-model approaches.
- GR-1 (ByteDance)字节 GR-1A late-2023 ByteDance manipulation model that first learns to predict future video frames, then learns to output actions.
- GR-2 (ByteDance)字节 GR-2ByteDance's second-generation manipulation model, released in 2024, pretrained first on 38 million web videos before learning to act.
- Seed GR-3字节 GR-3ByteDance Seed's 2025 general-purpose robot VLA, about 4 billion parameters, controlling its own ByteMini bimanual robot.
- GR-Dexter字节 GR-DexterA ByteDance Seed VLA from late 2025 for bimanual dexterous hands, built together with its own hand hardware and teleoperation system.
- GR-RL字节 GR-RLByteDance Seed uses reinforcement learning to turn a generalist VLA into a specialist, the first learned policy to lace a shoe autonomously.
- Robix字节 RobixByteDance Seed's high-level robot 'brain' model that unifies conversation, reasoning, and task planning.
- ERA-42星动纪元 ERA-42RobotEra's end-to-end VLA embodied model, released in late 2024, driving the company's dexterous hands and humanoid robots.
- AgiBot GO-1智元 GO-1(启元大模型)AgiBot's 2025 general-purpose embodied foundation model, using latent actions so human video can help train it too.
- AgiBot GO-2智元 GO-2(Genie Operator-2)AgiBot's 2026 second-generation embodied foundation model: it first plans a coarse action sequence, then executes it.
- A generative robot model from AgiBot and others that first predicts future multi-view frames with video diffusion, then produces actions.
- Genie Envisioner (AgiBot)智元 Genie Envisioner 世界模型AgiBot's platform that unifies a video world model, a policy, a neural simulator, and evaluation into one system.
- RoboBrain智源 RoboBrain(具身大脑)An open-source series of embodied 'brain' models from BAAI in Beijing, responsible for task planning and spatial understanding.
- GOVLA (Global & Omni-body Vision-Language-Action)智平方 GOVLA(AlphaBrain)AI² Robotics' large VLA model that outputs both full-body actions and a movement trajectory at once, across arbitrary embodiments and spaces.
- GraspVLA银河通用 GraspVLAA 2025 grasping foundation model from Galbot and others, pretrained mainly on a billion frames of simulated synthetic data.
- AstraBrain银河星脑 AstraBrainGalbot's embodied model family, connecting a “brain” that plans with a “cerebellum” that handles real-time whole-body control.
- Alibaba DAMO Academy's model that merges a VLA and a world model into one autoregressive system, outputting actions and predicting the next frame.
- RynnVLA-002达摩院 RynnVLA-002An open-source VLA from Alibaba DAMO Academy that merges an action model and a world model into one autoregressive network.
- RynnBrain达摩院 RynnBrainAlibaba DAMO Academy's open-source embodied 'brain' model, for understanding first-person video, localizing objects, and planning tasks.
- Being-H0智在无界 Being-H0BeingBeyond's series of dexterous-manipulation VLA models, pretrained on large-scale video of human hands doing manipulation.
- Galaxea's 2025 open-source fast-slow dual-system robot model, where a VLM breaks down the task and a VLA executes the actions.
- EO-1EO-1(EmbodiedOneVision)A 3B open-source unified embodied model from Shanghai AI Lab where a single network both reasons in language and outputs actions.
- InternVLA (Shanghai AI Laboratory)上海AI实验室 InternVLA 系列A family of embodied models from Shanghai AI Lab: M1 and A1 for manipulation, N1 for navigation.
- WALL-A自变量 WALL-AX Square Robot's in-house, closed-source embodied manipulation model series, running end-to-end from perception to motor control.
- WALL-OSS自变量 WALL-OSSX Square Robot's open-source embodied foundation model that turns a vision-language model into a VLA that outputs actions directly.
- UnifoLM宇树 UnifoLM 系列Unitree's open-source series of robot foundation models, spanning a world model, a VLA, and a general-purpose humanoid model.
- WoWWoW 具身世界模型A 14-billion-parameter embodied video world model trained on 2 million real-robot interaction trajectories.
- Pelican-VL北京人形 Pelican-VLAn open-source vision-language 'brain' model for embodied robots, released by China's X-Humanoid center.
- XR-1北京人形 XR-1A VLA foundation model from X-Humanoid that pretrains across robot embodiments using a 'unified vision-motion encoding.'
- GigaBrain-0极佳 GigaBrain-0A VLA model series from GigaAI trained mainly on data generated by a world model.
- GigaWorld-0极佳 GigaWorld-0GigaAI's framework that uses a world model as a “data engine” to mass-produce robot training data.
- MiMo-Embodied (Xiaomi)小米 MiMo-EmbodiedAn open-source 7B vision-language model from Xiaomi covering both autonomous-driving and embodied-AI understanding and planning.
- Xiaomi-Robotics-0小米 Xiaomi-Robotics-0Xiaomi's open-source, 4.7-billion-parameter VLA, built to execute actions in real time and smoothly on a consumer GPU.
- Kairos (ACE Robotics)大晓机器人 开悟世界模型A 4-billion-parameter embodied world model from ACE Robotics that does understanding, video generation, and action prediction all at once.
- UBTech Thinker优必选 ThinkerUBTech's in-house embodied model stack, made up of a foundation model, a world model, and an action model.
- Spirit v1.5千寻 Spirit v1.5An open-source VLA foundation model from Spirit AI, built on the premise that messy, unscripted data makes a better pretraining set.
- LingBot-VLA (Robbyant)蚂蚁灵波 LingBot-VLARobbyant's open-source VLA foundation model, pretrained on about 20,000 hours of bimanual real-robot data across 9 arm configurations.
- LingBot-VA (Robbyant)蚂蚁灵波 LingBot-VARobbyant's open-source video-action world model that predicts future frames and outputs robot actions at the same time.
- LingBot-World (Robbyant)蚂蚁灵波 LingBot-WorldRobbyant's open-source interactive world model that generates the next frame in real time from keyboard and camera commands.
- DM0原力灵机 DM0An “embodiment-native” VLA model open-sourced by Dexmal in 2026, pretrained from the start on a mix of driving and robot data.
- HoloBrain-0地平线 HoloBrain-0A VLA framework Horizon Robotics open-sourced in early 2026 that feeds camera parameters and robot structure into the model as priors.
- TARS Robotics AWE它石智航 AWETARS Robotics' general-purpose embodied foundation model, whose main selling point is training on large-scale human manipulation data.
- Psi-R2灵初 Psi-R2A world-action model from the Chinese startup PsiBot, pretrained on roughly 100,000 hours of human manipulation data.
- Tencent HY-Embodied (Hunyuan Embodied)腾讯 HY-Embodied(混元具身)An open-source embodied foundation model series from Tencent Robotics X and the Hunyuan team, including both an embodied VLM and a VLA.
- Qwen-Robot Series千问 Qwen-Robot 系列Alibaba's Qwen team released this trio of 2026 embodied models, one each for manipulation, navigation, and world modeling.
- MiniCPM-Robot series (ModelBest / OpenBMB)面壁 MiniCPM-Robot 系列MiniCPM's embodied model family, focused on small, on-device models for manipulation and for following a target on command.
12.9World models and learning from video
Back to the world-model thread: training policies inside imagination first, then generating worlds and learning actions from video.
- World ModelsWorld Models 论文(Ha & Schmidhuber)A classic 2018 world-model paper that trained an agent's policy entirely inside a 'dream' the model learned on its own.
- A model-based reinforcement learning agent that learns a latent-space world model from pixels and plans actions by imagining outcomes.
- A DeepMind algorithm that masters Go and Atari through tree search inside a learned internal model, with no rules given.
- Runs the Dreamer world model directly on real robots; a quadruped learns to walk from scratch in an hour.
- A world-model reinforcement-learning algorithm that trains a policy by imagining rollouts, using one fixed set of hyperparameters.
- A 2025 Google DeepMind world-model agent that mined diamonds in Minecraft using only offline data.
- A reinforcement learning algorithm that plans inside a learned latent-space world model, using one set of hyperparameters across hundreds of control tasks.
- A world model that predicts the future in DINOv2's image-feature space, able to plan zero-shot toward new goals after training.
- Meta's self-supervised video model that predicts video in feature space, usable as a world model for planning robot actions.
- Stability AI's open-source image-to-video diffusion model, commonly used in robotics research as a backbone for video prediction.
- SoraSora(视频生成即世界模拟器)OpenAI's text-to-video model, introduced with a technical report arguing that video generation models can double as world simulators.
- Genie (Original)Genie(初代)Google DeepMind's model that learns a “playable world” from unlabeled video, a generative interactive environment.
- Google DeepMind's large-scale world model that generates a controllable 3D world from a single image.
- Google DeepMind's 2025 real-time interactive world model: one text prompt generates a virtual world you can walk through.
- A Gemini-based Google DeepMind agent that follows instructions, reasons, and self-improves across many different 3D game worlds.
- DIAMONDDIAMOND(扩散世界模型)Generates frames one at a time with a diffusion model as a world model, letting an RL agent train inside a generated game.
- GameNGenGameNGen(神经游戏引擎)A 2024 Google project that uses a diffusion model to generate playable DOOM footage in real time, replacing the game engine.
- Matrix-Game (Skywork)昆仑万维 Matrix-GameSkywork's open-source real-time interactive world model that generates game footage frame by frame from keyboard and mouse input.
- HunyuanWorld (Tencent)腾讯混元世界模型Tencent Hunyuan's open-source 3D world-generation series that turns text or an image into an explorable 3D scene.
- Marble (World Labs)Marble(World Labs)World Labs' world-generation product that turns text, images, or video into a persistent, exportable 3D scene.
- MANOMANO 手部模型The most widely used parametric 3D hand model, describing a hand with a small number of shape and pose parameters.
- VPTVPT(视频预训练)OpenAI's approach of first training an inverse dynamics model to label online video with actions, then learning to play Minecraft by imitation.
- An imitation-learning method that learns high-level plans from video of human hands playing freely, then low-level actions from a little teleoperation data.
- VRBVRB(从人类视频学可供性)A method that learns 'where to grasp and which way to move afterward' from human video and hands that affordance directly to a robot.
- ATMATM(任意点轨迹建模)Learns how any point in a scene will move from video, then uses that predicted motion to guide a robot policy.
- Pretrains a VLA on action-label-free video by first learning latent actions, then mapping them to real actions with little robot data.
- PhantomPhantom(无机器人训练)A method that trains robot policies purely from human demonstration videos by digitally replacing the human hand with a rendered robot arm.
- A framework that trains a cross-embodiment VLA by learning 'task-relevant latent actions' from video.
- A VLA pretrained on first-person human video that converts human hand motion into humanoid robot actions.
- A method that first generates a video of a task being completed from text, then infers the robot's actions from that video.
- A learned 'real-world simulator,' built from a video generation model, that responds to actions.
- AVDCAVDC(从无动作视频学动作)First generates a video of a robot doing the task, then uses optical flow to derive the actions — no action labels needed.
- A method that uses an image-editing diffusion model as a high-level planner, sketching a subgoal image for a low-level policy to reach.
- A manipulation method that first generates a video of a human performing the task, then has the robot execute it by following that video.
- Video Prediction Policy视频预测策略A robot policy that guides action generation using the internal 'prediction of the future' features inside a video diffusion model.
- SeerSeer(预测式逆动力学模型)An end-to-end manipulation policy that first predicts the upcoming frame, then infers the action from it using inverse dynamics.
- A robot model in which video and action share one latent representation, letting inference skip video generation for speed.
- UWMUWM(统一世界模型)A robot framework that merges action diffusion and video diffusion into one model and can pretrain on action-free video.
- Vidar生数 VidarA robot manipulation model that first predicts future frames with a video diffusion model, then infers the action from them.
- A manipulation world model that generates multi-view future frames from actions, used to evaluate and improve policies “in imagination.”
- A latent-action world model that combines understanding, video generation, and action into three experts inside one model.
- A 2026 Tsinghua and Galaxea world action model that learns video prediction during training but skips imagining the future at inference, making it faster.
- FLUX 3 ActionFLUX 3 Action(Black Forest Labs)A 7B open-weight world action model from Germany's Black Forest Labs, released September 2026, predicting frames and actions together.
12.10Legged, dexterous-hand, and agile skills
From manipulation to movement: quadrupeds doing parkour, dexterous hands spinning objects, trained with RL in sim, then deployed to hardware.
- ANYmal RL Locomotion SeriesANYmal 强化学习运控系列(执行器网络 / 教师-学生盲走 / 感知行走)Three Science Robotics papers (2019, 2020, 2022) from ETH Zurich that took simulation-trained RL locomotion out into the real wild on the ANYmal quadruped.
- Rapid Motor Adaptation快速运动适应A method that lets quadruped robots sense terrain or load changes and adjust gait within a fraction of a second.
- A quadruped locomotion controller that learns many gaits in one policy and can be retuned live at deployment to handle new terrain.
- A 2023 KAIST reinforcement-learning method for blind quadruped locomotion that “imagines” the terrain underfoot using only proprioception.
- A quadruped-robot agility benchmark from Google, modeled on dog agility trials, for measuring animal-level agility in a repeatable way.
- A parkour policy that lets a low-cost quadruped robot dog climb, jump, crawl, and squeeze through obstacles using only a depth camera.
- A 2023 CMU project where a low-cost quadruped does parkour — jumping high and far — using just one depth camera and a single network.
- HIMHIM(混合内部模型)A legged-locomotion method that uses only proprioception, inferring terrain and disturbances from the robot's own response.
- ABSABS(敏捷且安全的足式运动)CMU's 2024 quadruped framework: sprint at high speed, switching to a learned collision-avoidance policy whenever a learned safety value says to.
- DIAL-MPCDIAL-MPC(扩散式退火足式 MPC)A training-free, sampling-based MPC for legged robots that borrows diffusion-style annealing to optimize whole-body motion in real time.
- DactylDactyl(OpenAI 魔方灵巧手)OpenAI's project that trained a five-fingered robot hand entirely in simulation with reinforcement learning, then transferred it directly to the real hand.
- HORAHORA(手内物体旋转 + RMA)A reinforcement-learning method that uses only fingertip and joint proprioception to keep rotating a variety of objects in-hand.
- Visual DexterityVisual Dexterity(视觉手内重定向)MIT's system that uses a single depth camera to let a low-cost dexterous hand reorient unseen objects in the air in real time.
- NVIDIA's arm-and-hand dexterous grasping system, trained entirely in simulation, grasping and carrying objects continuously from depth images alone.
- Meta and Berkeley's low-level dexterous-hand controller that turns a human's rough teleoperation intent into precise finger motion.
- Google DeepMind Table Tennis RobotDeepMind 乒乓球机器人DeepMind's 2024 table-tennis robot, the first learning-based robot to reach amateur human level in real matches.
12.11Humanoid whole-body control and teleop
The same methods carried over to humanoids: from animated-character imitation and bipedal walking to whole-body teleop and motion tracking.
- Uses reinforcement learning to make a simulated character imitate motion-capture clips, learning physically realistic flips and martial arts.
- ASEASE(对抗技能嵌入)Berkeley and NVIDIA's 2022 method for learning a reusable latent space of skills from motion-capture clips.
- MDM (Motion Diffusion Model)MDM(人体动作扩散模型)A landmark method that generates a 3D human motion sequence from text or an action category using a diffusion model.
- A physics-based humanoid controller that tracks huge libraries of human motion in simulation and gets back up on its own after falling.
- An NVIDIA unified physics-based humanoid controller that treats every control mode as filling in missing motion.
- Meta MotivoMeta Motivo(人形行为基础模型)Meta's simulated-humanoid behavior foundation model that does motion tracking, reaching a target pose, and reward-driven behavior with no further training.
- A Transformer walking controller trained purely with reinforcement learning in simulation, deployed zero-shot to make a humanoid walk outdoors.
- OP3 Soccer (DeepMind)OP3 足球(DeepMind 双足踢球)DeepMind uses deep reinforcement learning to teach the small humanoid robot OP3 to play one-on-one soccer.
- Humanoid Locomotion as Next Token Prediction人形行走即下一个 token 预测Treats a humanoid robot's walking control as a language-model-style “predict the next token” problem, learned with a causal Transformer.
- An end-to-end parkour policy that lets a humanoid jump onto platforms, cross gaps, and clear hurdles using just its head depth camera.
- Denoising World Model LearningDWL(去噪世界模型学习)An end-to-end reinforcement-learning walking framework using only proprioception, letting a humanoid cross snow, slopes, and stairs.
- A whole-body controller where one policy lets a humanoid walk, run, jump, and hop, with adjustable step frequency and foot-lift height.
- HoST (Humanoid Standing-up)HoST(人形起身)A control framework that uses reinforcement learning to learn, from scratch, how a humanoid can stand up from many different postures.
- A 2024 UC San Diego method where a humanoid's upper body imitates human motion while the legs just track velocity steadily.
- A humanoid whole-body motion-tracking method from UC San Diego and others that lets the Unitree G1 walk, squat, and dance.
- H2OH2O(人到人形)A 2024 CMU humanoid teleoperation framework that uses a single RGB camera to make a robot mimic a person's whole-body motion in real time.
- CMU's 2024 humanoid whole-body teleoperation system: VR, cameras, voice, or GPT-4o can all drive the same robot.
- Stanford's 2024 humanoid system: one RGB camera lets a robot shadow a person in real time, then learn skills from it.
- A 2024 UT Austin and NVIDIA method that teaches a humanoid robot to manipulate objects from watching a single human video.
- A versatile neural whole-body controller from NVIDIA and others where one policy works across many different command modes.
- UH-1UH-1 / Humanoid-XA large model that learns from massive internet human videos and generates humanoid robot motion from text instructions.
- Uses real-robot data to train a correction model that closes the sim-to-real gap for agile humanoid moves.
- A 2025 Shanghai AI Lab humanoid teleoperation system that controls the whole body with an exoskeleton, gloves, and foot pedals.
- Stanford's 2025 whole-body teleoperation system for humanoids: the robot mirrors the operator's entire body in real time.
- A portable teleoperation system that collects whole-body humanoid data with a VR headset, with no motion-capture studio needed.
- A humanoid robot method that learns skills like climbing stairs and sitting down in a chair from ordinary phone-shot human videos.
- UCSD's 2025 humanoid whole-body control method, using trajectory optimization to help an RL policy bend and reach far.
- A 2025 CMU humanoid training framework that lets a robot push, pull, and carry heavy loads steadily while walking.
- A humanoid whole-body teleoperation system that tracks only the head and hands, correcting drift in closed loop.
- A humanoid whole-body motion-imitation framework that teaches a Unitree G1 highly dynamic moves like kung fu and dance.
- A dual-system framework that links a vision-language model to a humanoid whole-body controller through a “latent verb.”
- A whole-body control method that uses one unified policy to make a humanoid track many different human motions.
- A three-stage whole-body motion tracking framework that lets a single policy make a humanoid robot track a wide range of human motions.
- Trains a humanoid to reproduce human motion, then distills that into a guided diffusion model for new tasks.
- A humanoid motion tracker from Tsinghua, Peking University, and Galbot that keeps tracking a motion even when pushed, pulled, or loaded down.
- A humanoid whole-body control foundation model trained with no task reward at all, switching tasks through a simple “prompt.”
- NVIDIA's general-purpose humanoid whole-body control foundation model, built by scaling up motion tracking training.
12.12Navigation and autonomous driving
Finally, moving through the world: navigation from modular pipelines to foundation models, and driving from ALVINN to VLA.
- CLIP on WheelsCoW(CLIP on Wheels)A baseline that bolts an open-vocabulary model like CLIP onto a mobile robot to find objects from text, with no navigation training.
- Combines GPT-3, CLIP, and a visual navigation model so a robot can navigate by following natural-language instructions.
- A system that writes vision-language features into a 3D map, letting a robot navigate using phrases like 'three meters to the right of the chair.'
- GNMGNM(通用导航模型)A general visual-navigation model trained on mixed data from 6 robot types that can drive different robots.
- A visual navigation foundation model from Berkeley, trained on navigation data from many kinds of robots, that finds a goal from an image.
- A 2023 Berkeley navigation diffusion policy where one model can both explore freely and head toward a goal image.
- VLFMVLFM(视觉语言前沿地图)A method that scores exploration frontiers with a vision-language model to find a specified object in an unfamiliar environment, zero-shot.
- A 2024 video-LLM navigation method from Peking University and others that decides the next step using only monocular video.
- A Transformer navigation policy trained purely with large-scale on-policy reinforcement learning in simulation, then deployed straight to real robots.
- A hierarchical navigation system that has a long-context VLM watch a tour video, then follows an image-and-text instruction to find the destination.
- A late-2024 Meta video world model for navigation that imagines what you'd see after taking a given action.
- A legged-robot navigation VLA where a large model gives mid-level actions in text and an RL locomotion controller does the walking.
- TrackVLA银河通用 TrackVLAA Galbot VLA for embodied visual tracking that simultaneously recognizes and plans a path to follow a target, using only its own first-person camera.
- NavFoM (Galbot)银河通用 NavFoMA 2025 navigation foundation model from Galbot and Peking University, with one set of weights adapting to many embodiments and navigation tasks.
- Carnegie Mellon's 1988 driving neural network, which maps camera images of the road directly to a steering command.
- UniADUniAD(规划导向的端到端自动驾驶)An end-to-end autonomous-driving framework that chains perception, prediction, and planning into one network, with everything serving the final plan.
- Tesla FSD v12特斯拉 FSD V12(端到端自动驾驶)Tesla's 2024 driver-assist release that replaced its city-street driving stack with a single end-to-end neural network.
- DriveVLMDriveVLM(快慢双系统智驾)A 2024 Tsinghua and Li Auto autonomous-driving system where a vision-language model thinks slowly and a traditional planner executes fast.
- GAIA-2 (Wayve)Wayve GAIA-2A 2025 driving world model from UK self-driving company Wayve that generates controllable, multi-camera driving video.
- NVIDIA Alpamayo英伟达 Alpamayo(驾驶推理 VLA)NVIDIA's open autonomous-driving reasoning VLA, which writes out in text why it's driving this way before outputting a trajectory.
- Waymo World ModelWaymo 世界模型Waymo's Genie-3-based generative driving simulator, able to generate camera and lidar data together.