Landmark 2 stops to get here · leads to 9

Overfitting

When a model fits its training data too closely, noise included, so it scores well on examples it has seen and poorly on new ones.

Your route here

2 stops · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training Data ✓ understood

    The examples a model learns its weights from, kept separate from the validation and test data used to check how well it generalizes.

  3. Overfitting · you are here ✓ understood

Picture it

training stepslossbest checkpointtrainingvalidationoverfitting →
Notice training loss keeps falling while validation loss turns back up: past the marked checkpoint the model is memorizing, not learning.

Overfitting is when a model learns its training data too specifically: it fits the particular examples, noise and quirks included, instead of the general pattern behind them. The symptom is a gap. Performance on the training data keeps improving while performance on new data stalls or gets worse.

The textbook picture is fitting a curve to a handful of points that really follow a gentle parabola. A straight line can’t bend enough and misses the pattern; that’s underfitting. A parabola fits well. A ninth-degree polynomial can pass exactly through every point and swing wildly in between. Its training error is zero, and it’s useless for prediction.

How to spot it

The diagram shows the standard check. Hold back a validation set the model never trains on, and track the loss on both sets as training goes. Early on, both fall together. At some point the validation loss bottoms out and starts to climb while the training loss keeps dropping. Whatever the model learns after that point is specific to the training set. Early stopping simply keeps the checkpoint where validation loss was lowest.

Why it happens

It comes down to capacity versus data: a model flexible enough to represent many different functions, with too few examples to pin down the right one. Modern networks have enormous capacity. In a 2016 study, standard image networks reached zero training error on CIFAR-10 even after every label was replaced at random, while their test accuracy was, necessarily, no better than chance. They can memorize. What keeps them from doing so on real data is a mix of the data itself and how they’re trained.

What helps

  • More data, or data augmentation, which creates varied copies of the examples you already have.
  • Regularization: penalties that favour simpler solutions, such as weight decay.
  • Dropout: randomly switching off units during training so the network can’t lean on any single one.
  • Early stopping, and a smaller model when the data really is small.

The catch

A validation set only protects you while it stays unseen. Tune enough settings against the same validation set and you slowly overfit to it as well. That’s why a final test set is kept untouched until the very end, and why a score on a benchmark that thousands of models have been tuned against deserves some suspicion.

Where it sits

Explore nearby

In the research

All papers →

2 papers that build on Overfitting .

Sources

  1. Ian Goodfellow, Yoshua Bengio and Aaron Courville, Deep Learning . MIT Press, 2016; section 5.2, Capacity, Overfitting and Underfitting; chapter 7, Regularization
  2. Zhang et al., "Understanding deep learning requires rethinking generalization" . 2016; ICLR 2017
  3. Srivastava et al., "Dropout: A Simple Way to Prevent Neural Networks from Overfitting" . Journal of Machine Learning Research, 2014