Fréchet Video Distance
弗雷歇视频距离FVDAdvancedA metric for how close a batch of generated videos is to real videos overall; lower is better.
FVD was proposed in late 2018 by researchers at Johannes Kepler University, IDSIA, and Google Brain, as the video counterpart to the image metric FID (Fréchet Inception Distance). It works by feeding both real and generated videos through an I3D network (a 3D convolutional network that captures both spatial appearance and motion over time) pretrained on the Kinetics action-recognition dataset to extract features, fitting each set to a multivariate Gaussian distribution, and computing the Fréchet distance between the two distributions. A lower value means the generated videos are closer to the real ones in both image quality and temporal coherence; the authors' human evaluation showed it correlates reasonably well with subjective human judgment. It's one of the most commonly used metrics in video generation and world model papers, but because it compares overall distributions, it doesn't check whether the physics or robot actions in any individual video are correct, which is why embodied settings often pair it with dedicated benchmarks such as EWMBench.
ExampleTo evaluate a robot video world model, a batch of real manipulation videos and videos the model generates from the same starting frames are both run through I3D to extract features; a smaller resulting FVD means the generated videos look more like real ones overall.
- Also called
- FVD
- Related
- Fréchet Inception Distance · Video Generation Model · World Model · EWMBench · VBench: Comprehensive Benchmark Suite for Video Generative Models · Peak Signal-to-Noise Ratio / Structural Similarity Index / Learned Perceptual Image Patch Similarity
- Sources
- Towards Accurate Generative Models of Video: A New Metric & Challenges (arXiv 1812.01717)