Embodied AI Glossary中文

GAIA-2 (Wayve)

Wayve GAIA-2GAIAAdvanced

A 2025 driving world model from UK self-driving company Wayve that generates controllable, multi-camera driving video.

GAIA-2 is a generative world model released by the UK self-driving company Wayve on March 26, 2025, the successor to GAIA-1. “World model” here means a video-generation model that can “imagine” future driving-scene footage conditioned on given inputs. GAIA-2 moves away from GAIA-1's autoregressive token generation to a video-tokenizer-plus-latent-diffusion architecture (compressing video into a latent space first, then denoising and generating within that latent space), natively supporting several cameras generated at once that stay consistent with each other in both time and space, with data covering the UK, the US, and Germany. Generation can be controlled along several axes: the ego vehicle's speed and steering, the behavior of other vehicles and pedestrians, weather and time of day, and road structure such as lanes and intersections. It is used to mass-produce synthetic data, create variations on real driving logs, and generate rare, dangerous scenarios to test driving models. It is one of the earlier examples of world models being put to practical use in physical AI.

ExampleTake a real multi-camera driving log, change the weather to heavy rain, and add a car suddenly cutting in, to generate a new set of surround-view video used to check how a driving model reacts to this rare situation.

Also called
A Controllable Multi-View Generative World Model for Autonomous Driving
Related
World Model · Autonomous Driving · Latent Diffusion Model · Synthetic Data · Waymo World Model · NVIDIA Cosmos
Sources
GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving (arXiv 2503.20523)
GAIA-2(Wayve 官方博客) (Chinese)
As of
2025-03

See it in the full glossary →