Softmax turns a list of arbitrary scores into probabilities. Give it any real numbers, called logits, and it returns the same number of values, each between 0 and 1, adding up to 1. It’s the last step of almost every classifier, and of every language model choosing its next token.
The recipe has two moves: raise e to the power of each score, then divide each result by their total. The exponential makes every value positive; the division makes them sum to 1.
What the numbers do
In the diagram the logits are 2.0, 1.0, 0.1 and −1.0, and softmax turns them into roughly 64%, 24%, 10% and 3%. Two properties explain the shape:
- Only the gaps matter. Add the same constant to every logit and the output doesn’t change. Implementations use this to stay numerically safe, subtracting the largest logit before exponentiating.
- The exponential stretches gaps. Each extra unit of logit multiplies the probability by e, about 2.7. A is only one point ahead of B, yet gets 2.7 times its share. Softmax is a “soft” version of picking the maximum: the leader gets most, but never all.
Where it’s used
- Classification. A network’s last layer produces one logit per class, softmax turns them into probabilities, and training scores them with cross-entropy. The pairing is deliberate: the log in the loss undoes the exponential, so the model keeps getting a useful learning signal even when it’s badly wrong.
- Language models. One softmax over the entire vocabulary for every token generated.
- Attention. Inside a transformer, softmax turns similarity scores between tokens into weights that sum to 1.
With only two classes, softmax reduces to the sigmoid.
Temperature
Dividing the logits by a number called the temperature before softmax changes how peaked the result is. Above 1 flattens the distribution; below 1 sharpens it toward the top choice. Hinton and colleagues used high temperatures to expose a large model’s “soft” preferences when distilling it into a smaller one, and the same knob, softmax temperature, controls how adventurous a language model’s sampling is.
The catch
Softmax always produces a tidy distribution, even for inputs unlike anything the model has seen. The outputs rank the options the model knows relative to each other; they aren’t a guarantee of how often the model is right. Whether 90% really means right nine times in ten is a separate question, called calibration.