Object Detection
目标检测EssentialFinding what objects appear in an image and where, marking each with a bounding box.
Object detection is a foundational computer vision task: given an image, it outputs the category, bounding box (the rectangle enclosing the object), and confidence score for each object in it. An early representative was the Viola-Jones face detector, built on hand-crafted features; after deep learning took over, the R-CNN family of two-stage detectors emerged, followed by YOLO in 2015, which treats detection directly as a regression problem over boxes and class probabilities and can run in real time. Evaluation uses intersection over union (IoU, the overlap area between two boxes divided by their combined area) to judge whether a predicted box matches the ground truth, aggregated into mean average precision (mAP, the average of the average precision across all classes). Traditional detectors only recognize the classes they were trained on; open-vocabulary detectors such as Grounding DINO can find objects from an arbitrary text description instead. In robotics, detection is often the first step of a modular pipeline: find the box, then segment it, estimate its pose, and plan a grasp.
ExampleA user says ‘hand me the red cup.’ The system feeds ‘red cup’ into Grounding DINO to get a bounding box, passes that to SAM to cut out a mask, combines the mask with the depth map to compute the cup's 3D position, and hands it off to the arm to grasp.
- Also called
- Object Recognition and Localization
- Related
- Bounding Box · Open-Vocabulary Object Detection · YOLO · Grounding DINO · Intersection over Union · Instance Segmentation
- Sources
- Wikipedia: Object detection
You Only Look Once: Unified, Real-Time Object Detection (arXiv 1506.02640)
Grounding DINO (arXiv 2303.05499)