What each split shows in a four-way classifier check
Case context: A team is building a classifier from one data source, but the data they ultimately care about comes from a different source. They keep four splits: a fit set, a holdout from the same source as the fit set, a development set, and a final test set.
Question: In this setup, what does checking the model on the fit set, the same-source holdout, and the development/test data tell you?
Sample answer: Performance on the fit set tells you the model's training error. Performance on the holdout from the same source tells you how well the model handles new examples that still match the training distribution. Performance on the development or test data tells you how well the model is likely to do on the real problem it will face after deployment.
Key points:
- Fit set evaluation measures training error.
- Same-source holdout evaluation measures generalization to new data from the training distribution.
- Development/test evaluation measures performance on the deployment task.
Rubric: The response must correctly identify that: 1. Fit set evaluation measures training error. 2. Same-source holdout evaluation measures generalization to new data from the training distribution. 3. Development/test evaluation measures performance on the deployment task.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Which dataset best measures how well a model handles new examples from the same distribution as its training data?
Training error is computed by evaluating the model on the same examples used to fit it.
Which Set Measures the Performance You Care About?
Match each dataset in the four-way evaluation setup to what it is mainly used to measure.
Order the evaluations used to move from training fit to the final assessment of performance on the target distribution.
Which dataset or datasets should be used to estimate how well a model will perform on the distribution you ultimately care about?
A training-dev set should come from the same distribution as the dev and test sets so it measures performance in the target environment.
Training-Distribution Development Set
Connect Each Dataset to Its Evaluation Purpose
Using the Four-Dataset Framework to Isolate the Main Weakness
How do the different datasets in a four-set evaluation scheme serve different diagnostic purposes?
What each split shows in a four-way classifier check
What the training dev set tells you