Computer vision is the part of AI that turns pixels into useful facts. A photo arrives as a grid of numbers, three per pixel for red, green and blue. The job is to get from that grid to statements like “this X-ray shows a fracture,” “there’s a cyclist ahead in the left lane,” or “the total on this receipt is $42.10.”
For decades that meant hand-designing the visual cues: edge detectors, corner finders, colour histograms, with a classifier on top. Modern systems learn their own cues from labeled images, which is why the field today is mostly deep learning.
The main tasks
- Image classification: one label for the whole image.
- Object detection: a box and a label for each object.
- Semantic segmentation: a label for every pixel.
- Beyond those: reading text in images (OCR), estimating body pose, tracking motion across video frames, and generating images rather than analysing them.
The more precise the task, the more expensive the labels. Tagging a photo “cat” takes a second; tracing every object’s outline takes far longer, so segmentation datasets tend to be much smaller than classification ones.
How it got good
The turning point came in 2012. Krizhevsky, Sutskever and Hinton trained a large convolutional neural network on GPUs using ImageNet, a benchmark of over a million labeled photos in 1,000 categories, and beat the previous best results by a wide margin. Convolutional networks dominated the field for the rest of the decade.
In 2020 the Vision Transformer showed that a transformer reading an image as a sequence of small patches could do just as well on classification when pre-trained on enough data. Today’s multimodal models attach a vision component like this to a language model, so you can ask questions about a picture in plain words.
The catch
A vision model learns whatever correlations its training photos contain, including accidental ones. A model trained mostly on daytime street scenes can struggle at night; one that only saw cows in fields may miss a cow on a beach. High scores on a benchmark don’t guarantee the model works with your camera, your lighting and your edge cases, so it has to be tested on images from where it will actually run.