Training Jul 2012

Improving Neural Networks by Preventing Co-adaptation of Feature Detectors

Geoffrey E. Hinton et al. · arXiv preprint

arXiv:1207.0580

In short

During training, randomly switch off half the hidden units on every example. No unit can rely on specific partners, so each learns features that are useful on their own, and overfitting drops sharply on speech and vision benchmarks.

Why it matters

Dropout became the standard regularizer of the deep-learning era and is still used inside large models.

Read first

The 3 Field Guide ideas this paper leans on.

Starting from scratch? The full route 7 ideas · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training Data ✓ understood

    The examples a model learns its weights from, kept separate from the validation and test data used to check how well it generalizes.

  3. Overfitting · read first ✓ understood

    When a model fits its training data too closely, noise included, so it scores well on examples it has seen and poorly on new ones.

  4. Regularization · read first ✓ understood

    Techniques to prevent overfitting by adding constraints or penalties to the model (L1, L2, dropout, early stopping).

  5. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  6. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  7. Dropout · read first ✓ understood

    A regularization technique that randomly deactivates neurons during training to prevent overfitting and improve generalization.

Nearby papers

Summary in our own words; read the paper for the details. ← All papers