DETR
AdvancedAn end-to-end detector that frames object detection directly as “set prediction” using a Transformer.
DETR is an object detection model proposed by Nicolas Carion and colleagues at Facebook AI (now Meta), published at ECCV 2020. Earlier detectors like Faster R-CNN first lay down a huge number of anchor boxes (preset candidate boxes) and finally use non-maximum suppression to delete duplicates. DETR first extracts image features with a CNN, feeds them into a Transformer encoder-decoder, and has the decoder use a fixed number of learnable “object queries” to each output one box and class; during training, bipartite matching pairs predictions with ground truth one to one, so anchor boxes and NMS are no longer needed. It matches a well-tuned Faster R-CNN’s accuracy on COCO, but converges more slowly during training. Later work — Deformable DETR, Grounding DINO, DINO-X — continued down this path; in embodied AI, ACT decodes actions with learnable queries too, with code adapted from DETR.
ExampleThe original DETR (ResNet-50 backbone), trained for 500 epochs on the COCO 2017 validation set, reaches 42.0 AP, on par with a same-backbone Faster R-CNN while using about half the computation.
- Also called
- DEtection TRansformer, End-to-End Object Detection with Transformers
- Related
- Object Detection · Transformer · Bounding Box · Non-Maximum Suppression · Grounding DINO · Learnable Query
- Sources
- arXiv 2005.12872: End-to-End Object Detection with Transformers
ECCV 2020 paper page: End-to-End Object Detection with Transformers
facebookresearch/detr (GitHub)