Landmark 3 stops to get here · leads to 16

Computer Vision

The field of AI that gets computers to extract meaning from images and video: what is in them, where it is, and how it moves.

Your route here

3 stops · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  3. Deep Learning ✓ understood

    A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.

  4. Computer Vision · you are here ✓ understood

Picture it

Classification

  • One label per image
  • "This is a cat"

Object detection

  • Box + label per object
  • "Cat here, dog there"

Segmentation

  • A label for every pixel
  • Exact object outlines
Notice how each task asks a more precise question about where things are in the same image.

Computer vision is the part of AI that turns pixels into useful facts. A photo arrives as a grid of numbers, three per pixel for red, green and blue. The job is to get from that grid to statements like “this X-ray shows a fracture,” “there’s a cyclist ahead in the left lane,” or “the total on this receipt is $42.10.”

For decades that meant hand-designing the visual cues: edge detectors, corner finders, colour histograms, with a classifier on top. Modern systems learn their own cues from labeled images, which is why the field today is mostly deep learning.

The main tasks

  • Image classification: one label for the whole image.
  • Object detection: a box and a label for each object.
  • Semantic segmentation: a label for every pixel.
  • Beyond those: reading text in images (OCR), estimating body pose, tracking motion across video frames, and generating images rather than analysing them.

The more precise the task, the more expensive the labels. Tagging a photo “cat” takes a second; tracing every object’s outline takes far longer, so segmentation datasets tend to be much smaller than classification ones.

How it got good

The turning point came in 2012. Krizhevsky, Sutskever and Hinton trained a large convolutional neural network on GPUs using ImageNet, a benchmark of over a million labeled photos in 1,000 categories, and beat the previous best results by a wide margin. Convolutional networks dominated the field for the rest of the decade.

In 2020 the Vision Transformer showed that a transformer reading an image as a sequence of small patches could do just as well on classification when pre-trained on enough data. Today’s multimodal models attach a vision component like this to a language model, so you can ask questions about a picture in plain words.

The catch

A vision model learns whatever correlations its training photos contain, including accidental ones. A model trained mostly on daytime street scenes can struggle at night; one that only saw cows in fields may miss a cow on a beach. High scores on a benchmark don’t guarantee the model works with your camera, your lighting and your edge cases, so it has to be tested on images from where it will actually run.

Where it sits

Explore nearby

Sources

  1. Richard Szeliski, Computer Vision: Algorithms and Applications . 2nd edition, Springer, 2022
  2. Krizhevsky, Sutskever and Hinton, "ImageNet Classification with Deep Convolutional Neural Networks" . NeurIPS 2012 (AlexNet)
  3. Dosovitskiy et al., "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale" . 2020 (Vision Transformer)