Embodied AI Glossary中文

Bird’s-Eye View

鸟瞰图BEVAdvanced

A 2D top-down grid that fuses information from multiple sensors onto a single ground plane.

A bird’s-eye view (BEV) is a 2D grid representation laid over the ground plane, viewed from directly above, where each cell stores features or semantics for that location. Autonomous driving was the first field to use it at scale: Lift-Splat-Shoot (NVIDIA, ECCV 2020) predicts a depth distribution for every pixel and then “splats” image features onto a BEV grid; BEVFormer (ECCV 2022, Shanghai AI Lab and others) uses Transformer cross-attention to query BEV features from multiple camera streams and fuses in past frames. The benefit is that multiple cameras and lidar can all be aligned into one coordinate frame, so detection, segmentation, and path planning can happen directly on this map. The occupancy grids, elevation maps, and semantic maps used in robot navigation are also essentially BEV representations; because BEV compresses away height information, tasks that depend more on height, like tabletop manipulation, usually switch to point clouds or voxels instead.

ExampleA self-driving car converts images from its six surround-view cameras into a single BEV feature map centered on itself, covering tens of meters ahead, behind, and to each side; it draws boxes around vehicles and lane lines directly on this map before handing it to the planning module.

Also called
BEV, BEV Perception, BEV Representation
Related
Autonomous Driving · 3D Object Detection · Occupancy Network · Elevation Map · Multi-Sensor Fusion · Semantic Map
Sources
BEVFormer (arXiv 2203.17270, ECCV 2022)
Lift, Splat, Shoot (arXiv 2008.05711, ECCV 2020)

See it in the full glossary →