Embodied AI Glossary中文

4D World Model

4D 世界模型Advanced

A world model that predicts how a 3D scene changes over time — 3D space plus time.

A world model predicts what the world will look like next, given the current observation and an action. Most video world models only generate 2D frames, with no depth or geometry, making it hard for a robot to read an object's precise position in space from them. A 4D world model instead predicts '3D space plus time': every frame carries geometric information, and the frames can be assembled into a 3D scene that evolves over time. The representative work is TesserAct, released in April 2025 by a UMass Amherst-led team: it fine-tunes the CogVideoX video generation model to jointly predict RGB, depth, and normal-vector video (RGB-DN), then reconstructs it into a temporally consistent 4D scene; actions are then computed by an inverse dynamics model (a network that infers the action from before-and-after states) from point-cloud features, producing 7-DOF robot-arm actions. It overlaps with 4D reconstruction, video generation, and spatial intelligence.

ExampleGiven the current view and a language instruction, TesserAct generates a future video with depth and normals, converts it into a 4D scene, encodes the resulting point cloud with PointNet, combines it with the instruction, and outputs a 7-DOF action.

Also called
4D Embodied World Model
Related
World Model · 4D Reconstruction · Video Generation Model · Inverse Dynamics Model · 3D VLA · Spatial Intelligence
Sources
TesserAct: Learning 4D Embodied World Models (arXiv 2504.20995)
TesserAct 论文 HTML 版 (Chinese)
As of
2025-04

See it in the full glossary →