Landmark 3 stops to get here · leads to 3
Regularization
Techniques to prevent overfitting by adding constraints or penalties to the model (L1, L2, dropout, early stopping).
Your route here
3 stops · basics first
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training Data ✓ understood
The examples a model learns its weights from, kept separate from the validation and test data used to check how well it generalizes.
- Overfitting ✓ understood
When a model fits its training data too closely, noise included, so it scores well on examples it has seen and poorly on new ones.
- Regularization · you are here ✓ understood
Picture it
Penalize weights
- L2 adds λ·Σw² to the loss
- L1 adds λ·Σ|w|, pushing weights to zero
- Weight decay shrinks weights each step
Inject noise
- Dropout zeroes random units in training
- Data augmentation varies the inputs
Stop early
- Watch validation loss during training
- Keep the checkpoint where it bottoms out
Where it sits
Explore nearby
Neural Networks Dropout A regularization technique that randomly deactivates neurons during training to prevent overfitting and improve generalization. Training Weight Decay A regularization technique that shrinks weights toward zero during optimization. Equivalent to L2 regularization in standard SGD, but differs when using adaptive optimizers like Adam. Training L1 Regularization Adding the sum of absolute weights to the loss function, promoting sparsity and feature selection. Training L2 Regularization Adding the sum of squared weights to the loss function, penalizing large weights and improving generalization. Training Early Stopping Stopping training when validation performance stops improving, preventing overfitting. Training Data Augmentation Creating variations of training data through transformations (rotation, cropping, noise) to improve model generalization.
In the research
All papers →A paper that builds on Regularization .