Stable Video Diffusion
SVDAdvancedStability AI's open-source image-to-video diffusion model, commonly used in robotics research as a backbone for video prediction.
Stable Video Diffusion is an open-source video generation model released by Stability AI on November 21, 2023, a latent diffusion model (which compresses images into a low-dimensional latent space and denoises step by step within it). It adds temporal layers on top of the Stable Diffusion image model and summarizes training into three stages — text-to-image pretraining, large-scale video pretraining, and high-quality video fine-tuning — emphasizing how much video-data curation matters. The initial release included two image-to-video models generating 14 and 25 frames respectively (the latter called SVD-XT), with a settable frame rate of 3 to 30 fps; it launched as a research preview, not for commercial use. Thanks to its open weights and modest size of about 1.5 billion parameters, SVD became a commonly used video backbone in embodied AI — for example, Video Prediction Policy (VPP) adds language conditioning on top of SVD, fine-tunes it on robot video, and extracts predictive representations of the future from it to output actions.
ExampleVideo Prediction Policy (VPP) fine-tunes SVD into a manipulation-video prediction model: given the current frame and the instruction 'open the drawer,' it first predicts features of the upcoming frames and then uses them to output the robot arm's actions.
- Also called
- SVD-XT, Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- Related
- Video Generation Model · Latent Diffusion Model · Video Prediction Policy · Text-to-Video / Image-to-Video · World Model · Diffusion Model
- Sources
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets (arXiv 2311.15127)
Introducing Stable Video Diffusion (Stability AI)
Video Prediction Policy (arXiv 2412.14803) - As of
- 2023-11