Embodied AI Glossary中文

Video Tokenizer

视频分词器Advanced

A codec network that compresses video into a small number of tokens or latent vectors and can reconstruct the footage from them.

A video tokenizer is a front-end module used by video generation models and world models. An encoder compresses a video clip across both time and space at once into either discrete tokens or continuous latent vectors, and a decoder reconstructs pixels from them. Models that output discrete tokens (such as Google's MAGVIT-v2) quantize vectors into “vocabulary indices” that plug neatly into an autoregressive Transformer; models that output continuous vectors are video VAEs (variational autoencoders), which a latent diffusion model then denoises inside. “Causal” means that along the time axis each frame only looks at earlier frames, with the first frame encoded on its own — this lets images and video share one tokenizer and lets arbitrarily long videos be processed in segments. The compression ratio determines how many tokens the downstream large model must process, while reconstruction quality caps how sharp the generated video can look. NVIDIA's Cosmos Tokenizer offers continuous and discrete versions at several compression ratios, including 4×8×8 and 8×16×16.

ExampleAlibaba's Wan2.1 uses a 3D causal VAE called Wan-VAE that compresses time 4× and each spatial dimension 8×; with feature caching it can encode and decode 1080p video of any length. The diffusion Transformer only denoises inside this compressed latent space, and the decoder reconstructs the final footage.

Also called
Video VAE, 3D Causal VAE, Causal Video VAE
Related
Variational Autoencoder · Latent Space · Spacetime Patches · Vector Quantization · Latent Diffusion Model · NVIDIA Cosmos
Sources
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation (MAGVIT-v2, arXiv:2310.05737)
NVIDIA Cosmos Tokenizer (GitHub)
Wan: Open and Advanced Large-Scale Video Generative Models (arXiv:2503.20314)

See it in the full glossary →