Supervised learning is learning from examples that come with the answer attached. Each training example is a pair: an input x (an email, a photo, a house’s size and location) and the correct output y (spam, “cat”, its sale price). The model’s job is to learn a mapping from x to y that also works on inputs it has never seen.
The name comes from the idea of a teacher. The label tells the system what it should have said, every time, and all of its learning comes from comparing its guesses against those answers.
How it works
The loop in the diagram runs thousands or millions of times. The model makes a prediction, a loss function turns the gap between prediction and label into a single number, and an optimizer such as gradient descent nudges the parameters to make that number smaller. The goal isn’t to fit the training examples perfectly. It’s to do well on fresh ones, so progress is checked on a held-out validation set and training stops when that score stops improving.
Two kinds of answer
- Classification: y is a category. Spam or not, which digit, which of a thousand object types. The model usually outputs a probability for each class.
- Regression: y is a number. A price, a temperature, a delivery time.
Many richer tasks are combinations of the two. An object detector, for example, classifies each object and regresses the four coordinates of its box.
Where the labels come from
Labels are the expensive part. Someone has to mark each email, draw each box, transcribe each audio clip, and experts are needed for things like medical scans. Collecting labeled data often costs more than training the model.
That cost is why the neighbouring approaches exist. Unsupervised learning uses no labels at all. Self-supervised learning manufactures labels from raw data, for instance by hiding the next word and asking the model to predict it; that is how large language models are pre-trained. Their later fine-tuning on example conversations is supervised learning again.
The catch
A supervised model can only be as good as its labels. Inconsistent or biased labels are learned just as faithfully as correct ones. It also learns the relationship between x and y as it held in the training data. If the world shifts, say fraudsters change tactics, the model keeps confidently applying last year’s pattern until it is retrained on new labelled examples.