Learn Before
Choose Dev and Test Sets to Match the Main Goal When Feasible
After the development and test sets are defined, model selection will usually be driven by dev-set results, so the dev set should represent the real task the team most wants to optimize. If dev and test data come from different sources, a model may look strong on dev data yet fail on test data, and the reason for that gap may be unclear. When both sets are drawn from the same distribution, that same gap has a much cleaner interpretation: the model is fitting the dev set too closely, and the next step is often to gather more dev-set examples. If the two sets are mismatched, the test gap could also mean the test task is harder or simply different, which makes diagnosis much less certain.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Why Dev and Test Sets Matter
How to Select Dev and Test Data for the Future Task
Choose Dev and Test Sets to Match the Main Goal When Feasible
Revise Evaluation Data and Metrics When They Stop Serving the Goal
Set the First Development and Test Splits Early for a New Project
Choosing a Dev Set Large Enough to See Small Gains
What is the main role of a development set during model iteration?
Another Name for the Development Set
Another name for the development set
Match each development-split role to its description.
Order the steps for choosing between two candidate spam filters using a validation set.
Which option is NOT a typical purpose of a development set?
The validation set is used to update model weights in the same way as the training set.
The validation set is used to tune parameters, _____ features, and make other choices about the learning method.
Match each development-set use to the task that illustrates it.
Arrange the main stages of a project that uses a development set.
What is the development set used for in model building?
Choosing the right split for model selection decisions.
What are the main roles of a development set?
Learn After
Make the Dev Set Match the Main Improvement Goal
A Shared-Source Validation/Test Gap Usually Means Validation Overfitting
Why Different Dev and Test Distributions Complicate Diagnosis
Mismatched Validation and Test Splits Can Make Chance Matter More
If a model is tuned on a development set and then performs worse on a separate test set, even though both sets come from the same source distribution, what is the most likely explanation?
True or False: If validation and test data are drawn from different populations, a score gap between them always has one clear cause.
If a model has started fitting the validation set too closely and the training and validation data come from the same distribution, the usual remedy is to get more _____ data.
Why should the development set match the main goal of the project?
When validation and test data come from the same distribution, a strong validation score followed by a much weaker test score suggests overfitting to the validation set.
If the development set is being overused and the training and development data come from the same distribution, the practical fix is to collect more _____ data.
Match each development-and-test-set situation with its consequence for debugging.
Order the troubleshooting steps when validation performance is strong but holdout performance is weak.
Why a Test Set Can Score Lower Than a Dev Set
A test score from a different data distribution tells you exactly why the model failed.
After the development and test sets are set, the team will spend its effort improving _____ set performance.
Match each concept about dev and test set distributions to the correct description.
How to Choose Dev and Test Sets for Reliable Iteration
Explain what poor final performance means when the tuning set and final set do or do not match.
Why a model can look strong in development but weak in final evaluation
What is the diagnosis and remedy when test results lag far behind dev results under the same data distribution?