Visual Odometry
视觉里程计VOCommonCompares consecutive camera frames to estimate, frame by frame, how far and how much a camera has turned.
Visual odometry uses a sequence of images from one or more cameras to estimate the camera’s relative motion frame by frame, then accumulates those estimates into a trajectory. The name was coined by Nistér and colleagues in 2004, borrowed from wheel odometry, which accumulates wheel rotations to estimate displacement; unlike wheel odometry, it isn’t thrown off by wheel slip, and NASA used it on both of its Mars rovers. There are two main approaches: feature-based methods extract and match features across frames, while direct methods solve for motion from pixel brightness error. A single, monocular camera can only recover a trajectory with unknown scale; stereo cameras or adding an IMU (making it visual-inertial odometry) give true scale. Visual odometry only tracks motion locally between neighboring frames, so error accumulates into drift over time; adding loop closure and global optimization turns it into visual SLAM.
ExampleAs a quadruped robot walks down a corridor, its front-facing stereo camera matches feature points frame by frame, estimating how far forward and how many degrees it turned relative to the previous frame and accumulating this into a trajectory; after going all the way around a large loop back to the start, the estimated endpoint usually doesn’t line up with the actual starting point — that mismatch is drift.
- Also called
- VO
- Related
- Visual-Inertial Odometry · Visual SLAM · Simultaneous Localization and Mapping · Loop Closure Detection · Feature Points · Wheel Odometry
- Sources
- Scaramuzza & Fraundorfer, Visual Odometry Part I: The First 30 Years and Fundamentals (IEEE RAM 2011)
Wikipedia: Visual odometry