Landmark 5 stops to get here · leads to 8

Object Detection

Finding every object of interest in an image and giving each a class label, a confidence score and a bounding box.

Your route here

5 stops · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  3. Deep Learning ✓ understood

    A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.

  4. Computer Vision ✓ understood

    The field of AI that gets computers to extract meaning from images and video: what is in them, where it is, and how it moves.

  5. Bounding Box ✓ understood

    A rectangular box defined by coordinates that localizes an object in an image, used in object detection.

  6. Object Detection · you are here ✓ understood

Picture it

  1. 01 Input image
  2. 02 Backbone network A CNN or transformer extracts features
  3. 03 Candidate boxes Each with box coordinates + class scores
  4. 04 Non-max suppression Drops overlapping duplicates using IoU
  5. 05 Labeled boxes e.g. dog 0.94, bicycle 0.88
Notice detection answers two questions at once for every object: what it is (a class label) and where it is (a bounding box).

Object detection finds every object of interest in an image and says what each one is and where it is. The output is a list: a class label, a confidence score and a bounding box for each object found. That’s harder than image classification, which gives one label for the whole picture, because the number of objects isn’t known in advance and each needs its own box.

It’s the perception step behind driver-assistance systems, warehouse robots, shelf-scanning cameras and photo search.

How it works

Following the diagram: a backbone network turns the image into feature maps, a grid of learned descriptions of what’s where. A detection head then proposes many candidate boxes, each with coordinates and a score for every class. The network deliberately over-proposes, so a single dog is usually covered by several overlapping boxes. Non-max suppression cleans this up: keep the highest-scoring box, drop any others that overlap it heavily as measured by IoU, and repeat.

Three families

  • Two-stage detectors first propose regions likely to contain something, then classify and refine each one. Faster R-CNN (2015) built the proposal step into the network itself and ran at about 5 frames per second on a GPU with a very deep backbone.
  • One-stage detectors predict boxes and classes in a single pass. YOLO (2015) splits the image into a grid where each cell predicts boxes directly. It ran at 45 frames per second, trading some precision in box placement for speed.
  • Transformer detectors such as DETR (2020) predict a fixed-size set of objects and match them one-to-one with the true objects during training. That removes hand-designed steps like non-max suppression and anchor boxes.

How it’s measured

A detection counts as correct when the class is right and the box overlaps the true box enough. A common threshold is an IoU of 0.5, and stricter benchmarks average over several thresholds. Results are summarised as average precision per class, then averaged across classes into mAP.

The catch

Detectors are only as good as their box annotations, and drawing boxes is slow, careful work. Small, partly hidden or unusual objects remain the common failures. A detector also knows only the classes it was trained on: anything else is either missed or forced into the nearest familiar label.

Where it sits

Explore nearby

In the research

All papers →

2 papers that build on Object Detection .

Sources

  1. Ren et al., "Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks" . 2015
  2. Redmon et al., "You Only Look Once: Unified, Real-Time Object Detection" . 2015 (YOLO)
  3. Carion et al., "End-to-End Object Detection with Transformers" . 2020 (DETR)