Training data is the set of examples a model learns from. Every weight in the finished model was set by fitting these examples, so in a real sense the model is a compressed summary of its training data. What’s in it, the model can learn. What’s missing, it can’t.
It’s one slice of a larger dataset. The rest is held back on purpose, because a model’s score on the examples it trained on says little about how it will do on new ones.
Three splits, three jobs
The split in the diagram exists to keep evaluation honest.
- Training set: the model sees it again and again and adjusts its weights to reduce the loss on it.
- Validation set: never trained on, but used to make choices such as the learning rate, the model size, or when to stop. A common rule of thumb carves it out of the training pool, roughly 80% for training and 20% for validation.
- Test set: untouched until the end and used for the final score. Once it’s used to make decisions, it stops being a fair test.
The splits must not overlap. When the same example, or a near copy, sits in both training and test data, scores look better than the model really is. That’s data leakage.
What it looks like
These are illustrative, but typical. A spam filter might learn from tens of thousands of emails, each labelled spam or not. A medical image model might have a few thousand scans, each labelled by a specialist. A large language model is pre-trained on a vast amount of web text, books and code with no human labels at all, because the next word supplies the answer, then fine-tuned on a far smaller, carefully written set of examples.
Quality beats quantity, up to a point
More data generally helps, but only if it resembles what the model will face. The common problems are:
- Label errors: the model learns mistakes as faithfully as truths.
- Skew and gaps: groups, languages or conditions that are under-represented get worse predictions.
- Shift over time: the world moves on after the data was collected, known as data drift.
That’s why researchers have proposed documenting datasets the way electronic components are documented, with a datasheet recording why the data was collected, what it contains and how it was gathered.
The catch
A model can’t know more than its data covers, and it can’t be fairer than its data is. Too few examples for a model’s size leads to overfitting: the model memorizes the training set instead of learning the pattern.