Landmark 2 stops to get here · leads to 1
Test Set
A final portion of data unseen during training and validation, used for unbiased evaluation of model performance.
Your route here
2 stops · basics first
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Train-Test Split ✓ understood
Dividing a dataset into separate portions for training the model and evaluating its performance on unseen data.
- Test Set · you are here ✓ understood
Picture it
- 01 Split the data Lock the test set away first
- 02 Train on training set Fit the model's parameters
- 03 Tune on validation set Pick hyperparameters and checkpoints
- 04 Evaluate on test set once An unbiased estimate on unseen data
Where it sits
Explore nearby
Evaluation Validation Set A portion of data held out from training, used to tune hyperparameters and monitor overfitting. Evaluation Cross-Validation A technique for assessing model performance by partitioning data into subsets, training on some and validating on others. Evaluation Data Leakage When information from outside the training data is used to create the model, leading to overly optimistic performance estimates. Evaluation Benchmark A standardized dataset and task used to compare model performance across different approaches (ImageNet, GLUE, SuperGLUE). Training Overfitting When a model fits its training data too closely, noise included, so it scores well on examples it has seen and poorly on new ones.