Training Feb 2015

Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Sergey Ioffe et al. · ICML 2015

arXiv:1502.03167

In short

Each layer’s inputs keep shifting as the layers before it learn, which forces small learning rates and careful initialization. Batch normalization rescales each layer’s inputs using statistics from the current mini-batch, as a built-in step of the network, so training tolerates much higher learning rates and reached the same image-classification accuracy in 14 times fewer steps.

Why it matters

It made very deep networks practical to train, and normalization layers became a standard part of the recipe.

Read first

The 3 Field Guide ideas this paper leans on.

Starting from scratch? The full route 8 ideas · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Neural Network · read first ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  3. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  4. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  5. Loss Function ✓ understood

    A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.

  6. Gradient Descent · read first ✓ understood

    An optimization method that repeatedly moves a model's parameters a small step in the direction that most reduces the loss.

  7. Batch Size ✓ understood

    The number of training examples processed together in one forward/backward pass.

  8. Batch Normalization · read first ✓ understood

    A technique that normalizes layer inputs to stabilize and accelerate training by reducing internal covariate shift.

Nearby papers

Summary in our own words; read the paper for the details. ← All papers