Standard 4 stops to get here · leads to 1
Stochastic Gradient Descent
A variant of gradient descent that updates parameters using gradients computed on a single random training example at a time (though often used to refer to mini-batch gradient descent).
Your route here
4 stops · basics first
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training ✓ understood
The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.
- Loss Function ✓ understood
A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.
- Gradient Descent ✓ understood
An optimization method that repeatedly moves a model's parameters a small step in the direction that most reduces the loss.
- Stochastic Gradient Descent · you are here ✓ understood
Where it sits
Explore nearby
Training Mini-Batch Gradient Descent Computing gradients on small batches of data, balancing SGD's noise with full-batch GD's stability. Training Batch Gradient Descent Computing gradients using the entire dataset, providing stable but slow updates. Training Momentum An optimization technique that accelerates gradient descent by accumulating past gradients, helping escape local minima. Training Adam Optimizer An adaptive learning rate optimization algorithm combining momentum and RMSprop, widely used for training neural networks. Training Batch Size The number of training examples processed together in one forward/backward pass.