Learn Before
Revise the evaluation setup when it stops matching the goal
If the original dev set, test set, or evaluation metric is no longer pointing the team toward the product objective, update it promptly. A warning sign is that the metric prefers one model while the team believes another model would serve users better. Common reasons include a development or test distribution that does not match the real operating environment, repeated tuning that causes overfitting to the dev set, or a metric that captures the wrong business priority. After the evaluation setup changes, tell the whole team so everyone optimizes the same target.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Single-Score Model Evaluation
An Early Dev/Test Setup Speeds Model Iteration
Revise the evaluation setup when it stops matching the goal
Error Analysis
Reviewing Both Missed and Correct Dev-Set Examples
After a development set and a test set have been chosen, what are they mainly used for?
True or False: A development set score can provide a quick check on whether a new idea is moving a model in the right direction.
Dev and test sets give a quick check on how well a team’s _____ is performing.
Match each dev/test-set action with its main effect.
Arrange the basic workflow for using development and test sets to improve a model.
Explain why a dev set, test set, and metric help a machine learning team work efficiently.
Diagnosing a project with no evaluation set
In one to three sentences, explain how a validation set can guide which model ideas deserve more work.
How does dev set performance help a team choose which ideas to keep improving?
What do teams usually try once dev and test sets are fixed?
What do teams typically do after they have set up dev and test sets?
True or False: With a fixed development set and a single evaluation metric, a team can quickly tell whether a new idea is helping by a little or by a lot.
A development set and a chosen evaluation metric help a team compare ideas quickly and see whether each change is making progress.
Learn After
When Evaluation Data Does Not Match Deployment Data
When Repeated Validation Checks Distort Model Selection
When the Metric Rewards the Wrong Goal
When should your validation setup be revised?
True or False: If your initial validation split or evaluation metric turns out to be poorly chosen, you cannot revise it without abandoning the project.
If your evaluation metric no longer reflects your main objective, what should you change?
What is the clearest sign that your dev/test set or evaluation metric may need revision?
If a validation set or metric turns out to be poorly matched to the real goal, the team should rebuild the whole project before making any changes.
What to revise when the evaluation no longer matches the goal
Match each reason a validation metric can mislead the team to the recommended remedy.
What should a team do when its evaluation setup stops matching its goal?
When Validation Data Does Not Match Deployment Data
After revising your dev/test sets or evaluation metric, updating the project documentation is enough; the team does not need to be told about the new direction.
What should be expanded after repeated tuning to the validation set?
Match each situation to the underlying problem category it illustrates.
Order the reasoning steps for deciding whether to replace an evaluation metric that no longer matches the product goal.
When validation results stop matching the best product choice
When Evaluation Scores and Product Needs Disagree
What should a team do after the development set stops guiding decisions?