Learn Before
Why reviewing only mistaken labels can distort validation results
Case context: A retail recommendation team audits its validation set after noticing several wrong predictions. They correct label errors only for the examples the model predicted incorrectly and conclude that the validation score is now unbiased.
Question: What is wrong with this approach, and what should the team do instead if it wants an unbiased evaluation set?
Sample answer: The process is biased because it changes labels selectively. Some examples were mislabeled even though the model predicted them correctly, so those errors remain hidden. A fair audit requires checking labels from both groups—mistaken predictions and correct predictions—or leaving the validation labels unchanged unless both groups can be reviewed.
Key points:
- Identify that selective relabeling creates evaluation bias.
- Explain that mislabeled examples among correct predictions are still missed.
- Recommend auditing a mix of correct and incorrect predictions, or not altering the set at all if a balanced audit is not possible.
Rubric: The answer must explain that correcting only mistaken predictions leaves hidden label errors and should recommend checking both correct and incorrect predictions, or not changing the validation set.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Reviewing Only Mistakes Can Skew a Dev Set
When Selective Relabeling Is Acceptable
What problem can occur if you correct only the dev-set examples your model got wrong?
True or False: If you relabel only the validation examples your system missed, the validation result remains an unbiased estimate.
Fixing labels only on the development examples your model misclassifies can introduce bias into evaluation because those examples are not selected at random.
How can selective relabeling skew a development-set score?
Editing Only the Wrongly Labeled Dev Examples Can Skew Your Evaluation
To avoid label-fix bias, review labels on _____ dev examples, not only the ones your model got wrong.
Match each review policy to its impact on validation-set fairness.
Put the label-review workflow in the right order to avoid bias.
What happens to the measured development accuracy if labels are corrected only for examples the model got wrong?
Checking only the dev examples your model got wrong is enough to guarantee an unbiased evaluation set.
Selective relabeling misses examples the classifier gets _____, so some bad labels are never examined.
Match each label-audit term with its meaning.
Put the steps in order to show why correcting only mislabeled mistakes can distort validation results.
Why does correcting only the labels of examples the model got wrong distort measured accuracy?
Why reviewing only mistaken labels can distort validation results
State the main risk of fixing only misclassified dev set labels.