Shipping AI Mar 2015

Distilling the Knowledge in a Neural Network

Geoffrey Hinton et al. · NeurIPS 2014 Deep Learning Workshop

arXiv:1503.02531

In short

A small model learns from a big one by matching its full, softened output probabilities, not just the right answer. Those “soft targets” carry what the big model knows about which wrong answers are nearly right, so the small model gets much of its accuracy at a fraction of the cost.

Why it matters

Distillation is how big models become fast, cheap ones you can actually deploy.

Read first

The 4 Field Guide ideas this paper leans on.

Starting from scratch? The full route 9 ideas · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  3. Activation Function ✓ understood

    A non-linear function applied to neuron outputs that introduces non-linearity, enabling networks to learn complex patterns.

  4. Softmax · read first ✓ understood

    A function that turns a list of scores (logits) into probabilities that are all positive and sum to 1; the standard output of classifiers and language models.

  5. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  6. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  7. Knowledge Distillation · read first ✓ understood

    Training a smaller 'student' model to mimic a larger 'teacher' model, transferring knowledge while reducing size.

  8. Teacher Model · read first ✓ understood

    The larger, more accurate model in knowledge distillation that guides student training.

  9. Student Model · read first ✓ understood

    The smaller model in knowledge distillation learning to mimic the teacher's behavior.

Nearby papers

Summary in our own words; read the paper for the details. ← All papers