Embodied AI Glossary中文

EnerVerse

Advanced

A generative robot model from AgiBot and others that first predicts future multi-view frames with video diffusion, then produces actions.

EnerVerse was released in January 2025 by AgiBot, Shanghai AI Lab, and other teams, and was later accepted to NeurIPS 2025. Its approach is to predict the future first and decide on actions second: an autoregressive video diffusion model generates future “embodied space” frames segment by segment based on an instruction, using sparse context memory to support long-horizon tasks. To represent 3D scenes, the authors propose Free Anchor Views, a multi-view video representation whose viewpoints are chosen flexibly per task. EnerVerse-A is a policy head attached after the generative model that converts the predicted 4D representation into robot actions; EnerVerse-D is a data engine that combines the generative model with 4D Gaussian splatting, used to automatically synthesize data and narrow the gap between simulation and reality.

ExampleGiven a manipulation instruction, the model first generates future video clips, from several viewpoints, of the robot completing the task, and the policy head then outputs actions based on them; the paper reports producing an 8-step action chunk in about 280 milliseconds.

Also called
EnerVerse-A, EnerVerse-D, Envisioning Embodied Future Space for Robotics Manipulation
Related
Video Prediction Model · World Model · Genie Envisioner (AgiBot) · 3D Gaussian Splatting · Multi-View · Video Prediction Policy
Sources
EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation (arXiv 2501.01895)
EnerVerse 项目页 (Chinese)
As of
2025-11

See it in the full glossary →