Overfitting is when a model learns its training data too specifically: it fits the particular examples, noise and quirks included, instead of the general pattern behind them. The symptom is a gap. Performance on the training data keeps improving while performance on new data stalls or gets worse.
The textbook picture is fitting a curve to a handful of points that really follow a gentle parabola. A straight line can’t bend enough and misses the pattern; that’s underfitting. A parabola fits well. A ninth-degree polynomial can pass exactly through every point and swing wildly in between. Its training error is zero, and it’s useless for prediction.
How to spot it
The diagram shows the standard check. Hold back a validation set the model never trains on, and track the loss on both sets as training goes. Early on, both fall together. At some point the validation loss bottoms out and starts to climb while the training loss keeps dropping. Whatever the model learns after that point is specific to the training set. Early stopping simply keeps the checkpoint where validation loss was lowest.
Why it happens
It comes down to capacity versus data: a model flexible enough to represent many different functions, with too few examples to pin down the right one. Modern networks have enormous capacity. In a 2016 study, standard image networks reached zero training error on CIFAR-10 even after every label was replaced at random, while their test accuracy was, necessarily, no better than chance. They can memorize. What keeps them from doing so on real data is a mix of the data itself and how they’re trained.
What helps
- More data, or data augmentation, which creates varied copies of the examples you already have.
- Regularization: penalties that favour simpler solutions, such as weight decay.
- Dropout: randomly switching off units during training so the network can’t lean on any single one.
- Early stopping, and a smaller model when the data really is small.
The catch
A validation set only protects you while it stays unseen. Tune enough settings against the same validation set and you slowly overfit to it as well. That’s why a final test set is kept untouched until the very end, and why a score on a benchmark that thousands of models have been tuned against deserves some suspicion.