Landmark 4 stops to get here · leads to 5
Learning Rate
A hyperparameter controlling the step size in gradient descent - too high causes instability, too low slows convergence.
Your route here
4 stops · basics first
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training ✓ understood
The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.
- Loss Function ✓ understood
A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.
- Gradient Descent ✓ understood
An optimization method that repeatedly moves a model's parameters a small step in the direction that most reduces the loss.
- Learning Rate · you are here ✓ understood
Picture it
Where it sits
Before this
Gradient Descent Learning Rate
Explore nearby
Training Learning Rate Schedule A strategy for adjusting the learning rate during training (decay, warm-up, cosine annealing) to improve convergence. Training Hyperparameter Configuration settings external to the model (learning rate, batch size) that must be set before training begins. Training Adam Optimizer An adaptive learning rate optimization algorithm combining momentum and RMSprop, widely used for training neural networks. Training Momentum An optimization technique that accelerates gradient descent by accumulating past gradients, helping escape local minima. Training Warmup Gradually increasing the learning rate at training start to stabilize optimization. Training Batch Size The number of training examples processed together in one forward/backward pass.
In the research
All papers →A paper that builds on Learning Rate .