A loss function scores how wrong a model’s prediction is, as a single number. Zero means a perfect prediction; bigger means worse. During training the model’s only goal is to make the average loss over its training examples as small as possible, so the loss function is, in effect, the definition of what “good” means for that model.
That one number is what makes learning mechanical. Gradient descent needs something to push downhill, and backpropagation works out how much each weight contributed to it. The last arrow in the diagram, gradients flowing back into the weights, exists only because the loss changes smoothly as the prediction changes.
The common ones
- Mean squared error: the average of squared differences between prediction and target, the default for predicting numbers. Squaring punishes big misses far more than small ones: off by 10 costs 100, off by 1 costs 1.
- Mean absolute error: the average of absolute differences. A single wild outlier sways it much less. Huber loss blends the two, acting like squared error for small misses and absolute error for large ones.
- Cross-entropy: the standard for classification. The model outputs probabilities, and the loss is the negative log of the probability it gave the correct answer. Give the right class 0.9 and the loss is about 0.1. Give it 0.3, as in the diagram, and the loss is about 1.2. Give it 0.01 and the loss is about 4.6. Confident wrong answers are punished hardest. Language models use this at every position: the loss measures how surprised the model was by the actual next token.
Choosing one
The choice follows from what you’re predicting. Most neural networks are trained by maximum likelihood, which for probability outputs means cross-entropy. It also pairs well with softmax outputs: squared or absolute error can give tiny gradients when those outputs saturate, while cross-entropy keeps learning moving.
Sometimes the loss is reshaped for the problem. Focal loss (2017) down-weights examples the model already gets right, so an object detector isn’t swamped by the huge number of easy background regions in every image.
The catch
The loss is a stand-in for what you actually care about, and the model optimizes the stand-in. Accuracy, fairness or usefulness only count if they’re reflected in it. Low training loss doesn’t guarantee a good model either: pushed far enough, it can mean the model has memorized its examples, which is overfitting. That’s why loss is also tracked on data the model never trains on.