Embodied AI Glossary中文

Spacetime Patches

时空块Advanced

Small cubes cut jointly across time and space from a video, each treated as one token for a Transformer.

A spacetime patch is the video version of 'cutting an image into patches.' A Vision Transformer cuts an image into fixed-size squares, each treated as a token; video adds a time dimension, so it's instead cut into small cubes spanning several frames and several pixels — Google's ViViT (ICCV 2021) calls this a tubelet, commonly extracted with a 3D convolution. OpenAI popularized the term 'spacetime patches' when it released Sora in February 2024: a video-compression network first compresses the video into latent space, and the latent representation is then cut into spacetime patches and handed to a diffusion Transformer to denoise. The benefit is that videos of different resolutions, durations, and aspect ratios can all become token sequences of varying length and be trained together, with no need to crop or rescale first. Most video generation models and video world models today turn video into tokens in a similar way.

ExampleIf a compressed latent video is cut into blocks of '2 frames × 2 × 2 cells,' each block is flattened and mapped by a linear layer into one token; landscape and portrait videos simply differ in how many blocks they're cut into and how they're arranged, and can be trained in the same batch.

Also called
Spacetime Latent Patches, Tubelet
Related
Vision Transformer · Diffusion Transformer · Video Tokenizer · Video Generation Model · Token · Sora
Sources
Sora: A Review on Background, Technology, Limitations, and Opportunities (arXiv:2402.17177)
ViViT: A Video Vision Transformer (arXiv:2103.15691)
Wikipedia: Sora (text-to-video model)
As of
2024-02

See it in the full glossary →