Object detection finds every object of interest in an image and says what each one is and where it is. The output is a list: a class label, a confidence score and a bounding box for each object found. That’s harder than image classification, which gives one label for the whole picture, because the number of objects isn’t known in advance and each needs its own box.
It’s the perception step behind driver-assistance systems, warehouse robots, shelf-scanning cameras and photo search.
How it works
Following the diagram: a backbone network turns the image into feature maps, a grid of learned descriptions of what’s where. A detection head then proposes many candidate boxes, each with coordinates and a score for every class. The network deliberately over-proposes, so a single dog is usually covered by several overlapping boxes. Non-max suppression cleans this up: keep the highest-scoring box, drop any others that overlap it heavily as measured by IoU, and repeat.
Three families
- Two-stage detectors first propose regions likely to contain something, then classify and refine each one. Faster R-CNN (2015) built the proposal step into the network itself and ran at about 5 frames per second on a GPU with a very deep backbone.
- One-stage detectors predict boxes and classes in a single pass. YOLO (2015) splits the image into a grid where each cell predicts boxes directly. It ran at 45 frames per second, trading some precision in box placement for speed.
- Transformer detectors such as DETR (2020) predict a fixed-size set of objects and match them one-to-one with the true objects during training. That removes hand-designed steps like non-max suppression and anchor boxes.
How it’s measured
A detection counts as correct when the class is right and the box overlaps the true box enough. A common threshold is an IoU of 0.5, and stricter benchmarks average over several thresholds. Results are summarised as average precision per class, then averaged across classes into mAP.
The catch
Detectors are only as good as their box annotations, and drawing boxes is slow, careful work. Small, partly hidden or unusual objects remain the common failures. A detector also knows only the classes it was trained on: anything else is either missed or forced into the nearest familiar label.