Adam: A Method for Stochastic Optimization
Diederik P. Kingma et al. · ICLR 2015
arXiv:1412.6980
In short
Adam gives every parameter its own step size, using running averages of its recent gradients and of their squares. It is simple, cheap in memory, copes with noisy and sparse gradients, and usually works well with little tuning.
Why it matters
Adam (and its descendant AdamW) is the default optimizer for training almost every modern network, LLMs included.
Read first
The 4 Field Guide ideas this paper leans on.
Starting from scratch? The full route 7 ideas · basics first
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training ✓ understood
The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.
- Loss Function ✓ understood
A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.
- Gradient Descent · read first ✓ understood
An optimization method that repeatedly moves a model's parameters a small step in the direction that most reduces the loss.
- Learning Rate · read first ✓ understood
A hyperparameter controlling the step size in gradient descent - too high causes instability, too low slows convergence.
- Momentum · read first ✓ understood
An optimization technique that accelerates gradient descent by accumulating past gradients, helping escape local minima.
- Adam Optimizer · read first ✓ understood
An adaptive learning rate optimization algorithm combining momentum and RMSprop, widely used for training neural networks.