Text-to-Video / Image-to-Video
文生视频 / 图生视频T2V / I2VCommonGenerating a video from a text description, or from an image plus text.
Text-to-video (T2V) and image-to-video (I2V) are the two most common ways of using a video generation model. In T2V, the model generates video from scratch given only a text description; in I2V, it's given an image (usually treated as the first frame) plus text, and animates the scene forward from there. Stability AI's Stable Video Diffusion and Alibaba's Wan both offer versions of each; the dominant approach is to run diffusion or flow-matching denoising step by step inside a latent space produced by a VAE. I2V is more useful for robotics: treating the robot's current camera view as the first frame and the task instruction as the text turns the generated video into a preview of what should happen next. UniPi uses text-guided video generation for planning, while NVIDIA's DreamZero adds action output directly on top of Wan2.1's 14B image-to-video model.
ExampleGiven a photo of a robot arm and a cup on a table, plus the instruction 'put the cup in the sink,' an image-to-video model generates a few seconds of video starting from that photo, showing the arm carrying out the task.
- Also called
- T2V, I2V, Text2Video, Image2Video
- Related
- Video Generation Model · Diffusion Model · Latent Diffusion Model · Variational Autoencoder · World Action Model · UniPi
- Sources
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets (arXiv 2311.15127)
Wan: Open and Advanced Large-Scale Video Generative Models (arXiv 2503.20314)
World Action Models are Zero-shot Policies (arXiv 2602.15922) - As of
- 2026-02