Standard 4 stops to get here · leads to 1

Weight Decay

A regularization technique that shrinks weights toward zero during optimization. Equivalent to L2 regularization in standard SGD, but differs when using adaptive optimizers like Adam.

Your route here

4 stops · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training Data ✓ understood

    The examples a model learns its weights from, kept separate from the validation and test data used to check how well it generalizes.

  3. Overfitting ✓ understood

    When a model fits its training data too closely, noise included, so it scores well on examples it has seen and poorly on new ones.

  4. Regularization ✓ understood

    Techniques to prevent overfitting by adding constraints or penalties to the model (L1, L2, dropout, early stopping).

  5. Weight Decay · you are here ✓ understood

Where it sits

Before this

Regularization
Weight Decay

Leads to

AdamW

Explore nearby