Standard 4 stops to get here · leads to 2
Momentum
An optimization technique that accelerates gradient descent by accumulating past gradients, helping escape local minima.
Your route here
4 stops · basics first
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training ✓ understood
The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.
- Loss Function ✓ understood
A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.
- Gradient Descent ✓ understood
An optimization method that repeatedly moves a model's parameters a small step in the direction that most reduces the loss.
- Momentum · you are here ✓ understood
Where it sits
Explore nearby
Training Nesterov Momentum A momentum variant that looks ahead before computing gradients, often converging faster. Training Stochastic Gradient Descent A variant of gradient descent that updates parameters using gradients computed on a single random training example at a time (though often used to refer to mini-batch gradient descent). Training Adam Optimizer An adaptive learning rate optimization algorithm combining momentum and RMSprop, widely used for training neural networks. Training RMSprop An optimizer using moving average of squared gradients to adapt learning rates, addressing AdaGrad's diminishing rates.
In the research
All papers →A paper that builds on Momentum .