Embodied AI Glossary中文

Novel View Synthesis

新视角合成NVSAdvanced

Generating an image of a scene from a camera viewpoint that was never actually photographed, using a handful of existing photos.

Novel view synthesis takes several images of a scene, usually with known camera poses, and renders what the scene would look like from a new camera position. Classic approaches reconstruct geometry with multi-view stereo and then apply texture; Neural Radiance Fields (NeRF), introduced in 2020, instead represent the scene with a neural network and render it through volume rendering, which substantially improved quality. 3D Gaussian Splatting, from 2023, represents the scene explicitly as a large number of 3D Gaussians and can render in real time. More recently, video generation models have also been used to generate novel views directly. In embodied AI, novel view synthesis is used to bring real-world scenes into simulation (real-to-sim) and to augment training data for policies with extra viewpoints.

ExampleDozens of phone photos taken while walking around an object are used to train a 3D Gaussian Splatting model, which can then render the object from any new angle.

Also called
NVS, New View Synthesis
Related
Neural Radiance Fields · 3D Gaussian Splatting · Multi-View Stereo · Real-to-Sim · Gaussian Splatting-based Simulation · Video Generation Model
Sources
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis (arXiv)

See it in the full glossary →