Embodied AI Glossary中文

YOLO

Common

A family of real-time object detectors that locate every object in an image with a single pass through the network.

YOLO was introduced by Joseph Redmon and colleagues in 2015 (published at CVPR 2016), framing object detection as a regression problem: the whole image passes through the network only once, producing the location and class of every detection box simultaneously; the original version ran at 45 frames per second, far faster than two-stage methods that first propose candidate boxes and then classify each one. YOLO has since grown into a large family released in relays by different teams; Ultralytics’ YOLOv5, YOLOv8, and YOLO11 are the most widely used. According to its documentation, YOLO26, released in January 2026, can optionally drop non-maximum suppression (NMS, the post-processing step that removes duplicate boxes) and supports segmentation, pose, and rotated-box tasks. Robots commonly use YOLO for fast 2D detection, then combine it with a depth map to get an object’s 3D position; for finding arbitrary categories by text, there’s the open-vocabulary version, YOLO-World.

ExampleDesktop sorting: a wrist-camera image is fed into YOLO to get a bounding box for “cup”; the depth pixels inside the box are back-projected into 3D as the target position for the robot arm’s grasp.

Also called
You Only Look Once, YOLOv8, YOLO11, YOLO26
Related
Object Detection · Bounding Box · Non-Maximum Suppression · YOLO-World · Mean Average Precision
Sources
You Only Look Once: Unified, Real-Time Object Detection (arXiv 1506.02640)
Ultralytics Docs: Models Supported by Ultralytics
As of
2026-01

See it in the full glossary →