Landmark 1 stop to get here · leads to 8

Training Data

The examples a model learns its weights from, kept separate from the validation and test data used to check how well it generalizes.

Your route here

1 stop · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training Data · you are here ✓ understood

Picture it

Training data

  • Usually the largest share
  • The model learns its weights from it
  • Seen many times across epochs

Validation data

  • Held out from training
  • Used to tune and catch overfitting

Test data

  • Held out until the end
  • Gives the final unbiased score
Notice that only the training data ever changes the model's weights; the other splits exist to check how well that learning generalizes.

Training data is the set of examples a model learns from. Every weight in the finished model was set by fitting these examples, so in a real sense the model is a compressed summary of its training data. What’s in it, the model can learn. What’s missing, it can’t.

It’s one slice of a larger dataset. The rest is held back on purpose, because a model’s score on the examples it trained on says little about how it will do on new ones.

Three splits, three jobs

The split in the diagram exists to keep evaluation honest.

  • Training set: the model sees it again and again and adjusts its weights to reduce the loss on it.
  • Validation set: never trained on, but used to make choices such as the learning rate, the model size, or when to stop. A common rule of thumb carves it out of the training pool, roughly 80% for training and 20% for validation.
  • Test set: untouched until the end and used for the final score. Once it’s used to make decisions, it stops being a fair test.

The splits must not overlap. When the same example, or a near copy, sits in both training and test data, scores look better than the model really is. That’s data leakage.

What it looks like

These are illustrative, but typical. A spam filter might learn from tens of thousands of emails, each labelled spam or not. A medical image model might have a few thousand scans, each labelled by a specialist. A large language model is pre-trained on a vast amount of web text, books and code with no human labels at all, because the next word supplies the answer, then fine-tuned on a far smaller, carefully written set of examples.

Quality beats quantity, up to a point

More data generally helps, but only if it resembles what the model will face. The common problems are:

  • Label errors: the model learns mistakes as faithfully as truths.
  • Skew and gaps: groups, languages or conditions that are under-represented get worse predictions.
  • Shift over time: the world moves on after the data was collected, known as data drift.

That’s why researchers have proposed documenting datasets the way electronic components are documented, with a datasheet recording why the data was collected, what it contains and how it was gathered.

The catch

A model can’t know more than its data covers, and it can’t be fairer than its data is. Too few examples for a model’s size leads to overfitting: the model memorizes the training set instead of learning the pattern.

Where it sits

Explore nearby

In the research

All papers →

A paper that builds on Training Data .

Sources

  1. Ian Goodfellow, Yoshua Bengio and Aaron Courville, Deep Learning . MIT Press, 2016; section 5.3.1, Validation Sets
  2. Gebru et al., "Datasheets for Datasets" . 2018