Landmark 2 stops to get here · leads to 9

Loss Function

A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.

Your route here

2 stops · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  3. Loss Function · you are here ✓ understood

Picture it

  1. 01 Model prediction e.g. predicts 0.3 for the true class
  2. 02 True value The label from the training data
  3. 03 Loss function e.g. cross-entropy or mean squared error
  4. 04 One number: the loss Bigger means more wrong
  5. 05 Gradients update weights Weights move to make the loss smaller
Notice how the loss function turns "how wrong was that?" into a single number that training can push downhill.

A loss function scores how wrong a model’s prediction is, as a single number. Zero means a perfect prediction; bigger means worse. During training the model’s only goal is to make the average loss over its training examples as small as possible, so the loss function is, in effect, the definition of what “good” means for that model.

That one number is what makes learning mechanical. Gradient descent needs something to push downhill, and backpropagation works out how much each weight contributed to it. The last arrow in the diagram, gradients flowing back into the weights, exists only because the loss changes smoothly as the prediction changes.

The common ones

  • Mean squared error: the average of squared differences between prediction and target, the default for predicting numbers. Squaring punishes big misses far more than small ones: off by 10 costs 100, off by 1 costs 1.
  • Mean absolute error: the average of absolute differences. A single wild outlier sways it much less. Huber loss blends the two, acting like squared error for small misses and absolute error for large ones.
  • Cross-entropy: the standard for classification. The model outputs probabilities, and the loss is the negative log of the probability it gave the correct answer. Give the right class 0.9 and the loss is about 0.1. Give it 0.3, as in the diagram, and the loss is about 1.2. Give it 0.01 and the loss is about 4.6. Confident wrong answers are punished hardest. Language models use this at every position: the loss measures how surprised the model was by the actual next token.

Choosing one

The choice follows from what you’re predicting. Most neural networks are trained by maximum likelihood, which for probability outputs means cross-entropy. It also pairs well with softmax outputs: squared or absolute error can give tiny gradients when those outputs saturate, while cross-entropy keeps learning moving.

Sometimes the loss is reshaped for the problem. Focal loss (2017) down-weights examples the model already gets right, so an object detector isn’t swamped by the huge number of easy background regions in every image.

The catch

The loss is a stand-in for what you actually care about, and the model optimizes the stand-in. Accuracy, fairness or usefulness only count if they’re reflected in it. Low training loss doesn’t guarantee a good model either: pushed far enough, it can mean the model has memorized its examples, which is overfitting. That’s why loss is also tracked on data the model never trains on.

Where it sits

Explore nearby

In the research

All papers →

A paper that builds on Loss Function .

Sources

  1. Ian Goodfellow, Yoshua Bengio and Aaron Courville, Deep Learning . MIT Press, 2016; section 6.2.1, Cost Functions
  2. Lin et al., "Focal Loss for Dense Object Detection" . 2017