Improving Neural Networks by Preventing Co-adaptation of Feature Detectors
Geoffrey E. Hinton et al. · arXiv preprint
arXiv:1207.0580
In short
During training, randomly switch off half the hidden units on every example. No unit can rely on specific partners, so each learns features that are useful on their own, and overfitting drops sharply on speech and vision benchmarks.
Why it matters
Dropout became the standard regularizer of the deep-learning era and is still used inside large models.
Read first
The 3 Field Guide ideas this paper leans on.
Starting from scratch? The full route 7 ideas · basics first
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training Data ✓ understood
The examples a model learns its weights from, kept separate from the validation and test data used to check how well it generalizes.
- Overfitting · read first ✓ understood
When a model fits its training data too closely, noise included, so it scores well on examples it has seen and poorly on new ones.
- Regularization · read first ✓ understood
Techniques to prevent overfitting by adding constraints or penalties to the model (L1, L2, dropout, early stopping).
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Neural Network ✓ understood
A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.
- Dropout · read first ✓ understood
A regularization technique that randomly deactivates neurons during training to prevent overfitting and improve generalization.