Landmark 2 stops to get here · leads to 10

Classification

A supervised learning task where the model assigns each input to one of a fixed set of categories, such as spam or not spam.

Your route here

2 stops · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Supervised Learning ✓ understood

    Learning from examples paired with the correct answer, so a model can predict answers for new inputs it hasn't seen.

  3. Classification · you are here ✓ understood

Picture it

Classification

  • Predicts a discrete label
  • e.g. spam or not spam
  • Output: class probabilities
  • Scored by accuracy, F1

Regression

  • Predicts a continuous number
  • e.g. a house price
  • Output: a real value
  • Scored by MSE, MAE
Both are supervised learning; notice that what differs is whether the answer is a category or a number.

Classification is predicting which bucket something belongs in. Is this email spam? Which of 1,000 object types is in this photo? Is this card payment fraudulent? Is this review positive, negative or neutral? The answer is always one of a fixed list of classes, and the model learns to choose from labeled examples, which makes classification a form of supervised learning. When the answer is a number instead of a category, the task is regression.

Most classifiers output more than a label. They produce a score for each class, usually turned into probabilities with softmax, and the predicted label is the most likely one. The probabilities are useful in their own right: a fraud system might block a payment at 99% and send one at 55% to a person to review.

The kinds

  • Binary: two classes, such as spam or not spam. The model outputs one probability and you choose the threshold.
  • Multi-class: exactly one of many classes, such as which handwritten digit or which species of bird.
  • Multi-label: any number of classes at once. A news article can be about both politics and sport.

Almost any model can classify: logistic regression, decision trees, support vector machines, and neural networks, including language models used for text classification. Each learns a decision boundary, the line or surface in the space of inputs that separates one class from another.

Measuring it honestly

Accuracy, the share of predictions that are right, is the obvious metric and often the wrong one. If 1 payment in 1,000 is fraud, a model that always answers “not fraud” is 99.9% accurate and useless. A confusion matrix breaks results down by class. Precision asks how many of the flagged cases were real; recall asks how many of the real cases were flagged. Moving the threshold trades one for the other, and the right balance depends on which mistake costs more.

The catch

Labels come from people, and people disagree, make mistakes and draw category lines differently, so a classifier inherits that noise. A classifier also has no “none of the above” unless you give it one. Shown something outside all its classes, it will still pick one, often confidently.

Where it sits

Explore nearby

In the research

All papers →

2 papers that build on Classification .

Sources

  1. Trevor Hastie, Robert Tibshirani and Jerome Friedman, The Elements of Statistical Learning . 2nd edition, Springer, 2009
  2. Ian Goodfellow, Yoshua Bengio and Aaron Courville, Deep Learning . MIT Press, 2016; section 5.1, Learning Algorithms