Training Oct 1986

Learning Representations by Back-Propagating Errors

David E. Rumelhart et al. · Nature

doi:10.1038/323533a0

In short

The paper shows how to compute, layer by layer from the output backwards, how much each weight in a multi-layer network contributed to the error, and nudge it accordingly. Trained this way, hidden units learn useful internal features on their own.

Why it matters

Backpropagation is how essentially every neural network, from LeNet to today’s LLMs, is trained.

Read first

The 4 Field Guide ideas this paper leans on.

Starting from scratch? The full route 7 ideas · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Neural Network · read first ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  3. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  4. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  5. Loss Function · read first ✓ understood

    A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.

  6. Gradient Descent · read first ✓ understood

    An optimization method that repeatedly moves a model's parameters a small step in the direction that most reduces the loss.

  7. Backpropagation · read first ✓ understood

    The algorithm for computing gradients of the loss with respect to network weights, enabling training through gradient descent.

Nearby papers

Summary in our own words; read the paper for the details. ← All papers