Computer Vision
计算机视觉CVEssentialThe field of making computers extract useful information from images and video and understand what a scene contains.
Computer vision studies how to make a computer automatically extract, analyze, and understand information from a single image or a video sequence, covering tasks such as classification, object detection, segmentation, tracking, pose estimation, and 3D reconstruction. The field got its start in AI labs in the late 1960s — in 1966, someone reportedly thought hooking a camera up to a computer and having it ‘describe what it sees’ would take an undergraduate just one summer project; the problem has instead occupied researchers for more than half a century since. Since deep learning took off, convolutional networks, vision transformers, and vision foundation models such as CLIP, DINOv2, and SAM have become the mainstream tools. In embodied AI, the camera is a robot's main channel for information about the outside world, and the vision encoders inside VLA models mostly reuse models pretrained by the computer vision community directly.
ExampleBefore a robot clears a table, it first uses object detection to locate a cup, segmentation to cut out the cup's outline, and then combines that with a depth map to compute the cup's 3D position — all of these steps belong to computer vision.
- Also called
- CV
- Related
- Object Detection · Semantic Segmentation · 3D Vision · Vision Foundation Model · Depth Estimation · Vision Encoder
- Sources
- Wikipedia: Computer vision
Richard Szeliski, Computer Vision: Algorithms and Applications (2nd ed., 2022)