4D World Model
4D 世界模型AdvancedA world model that predicts how a 3D scene changes over time — 3D space plus time.
A world model predicts what the world will look like next, given the current observation and an action. Most video world models only generate 2D frames, with no depth or geometry, making it hard for a robot to read an object's precise position in space from them. A 4D world model instead predicts '3D space plus time': every frame carries geometric information, and the frames can be assembled into a 3D scene that evolves over time. The representative work is TesserAct, released in April 2025 by a UMass Amherst-led team: it fine-tunes the CogVideoX video generation model to jointly predict RGB, depth, and normal-vector video (RGB-DN), then reconstructs it into a temporally consistent 4D scene; actions are then computed by an inverse dynamics model (a network that infers the action from before-and-after states) from point-cloud features, producing 7-DOF robot-arm actions. It overlaps with 4D reconstruction, video generation, and spatial intelligence.
ExampleGiven the current view and a language instruction, TesserAct generates a future video with depth and normals, converts it into a 4D scene, encodes the resulting point cloud with PointNet, combines it with the instruction, and outputs a 7-DOF action.
- Also called
- 4D Embodied World Model
- Related
- World Model · 4D Reconstruction · Video Generation Model · Inverse Dynamics Model · 3D VLA · Spatial Intelligence
- Sources
- TesserAct: Learning 4D Embodied World Models (arXiv 2504.20995)
TesserAct 论文 HTML 版 (Chinese) - As of
- 2025-04