Learn Before
Why does correcting only the labels of examples the model got wrong distort measured accuracy?
Question: A team reviews a validation set for an image classifier and fixes labels only for the cases the model predicted incorrectly. What happens to the reported accuracy, and why does it differ from the model’s true performance?
Sample answer: This process can only turn some counted mistakes into correct predictions. Because the team checks only examples already marked wrong by the current labels, it never examines the cases the model predicted correctly. As a result, the measured number of errors goes down while the number of counted correct predictions goes up, even though the evaluation set has not been fully cleaned. The reported accuracy therefore becomes too optimistic and can overstate how well the system really performs.
Key points:
- Only examples already counted as errors are inspected and relabeled.
- Any label problems on examples the model predicted correctly are left untouched.
- The recorded accuracy becomes artificially high, so the evaluation is biased upward.
Rubric: The response must explain: 1. only examples the model missed are checked, so the measured error rate can only decrease; 2. correctly predicted examples with bad labels are not corrected; 3. this selective relabeling makes the reported accuracy overly optimistic.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Reviewing Only Mistakes Can Skew a Dev Set
When Selective Relabeling Is Acceptable
What problem can occur if you correct only the dev-set examples your model got wrong?
True or False: If you relabel only the validation examples your system missed, the validation result remains an unbiased estimate.
Fixing labels only on the development examples your model misclassifies can introduce bias into evaluation because those examples are not selected at random.
How can selective relabeling skew a development-set score?
Editing Only the Wrongly Labeled Dev Examples Can Skew Your Evaluation
To avoid label-fix bias, review labels on _____ dev examples, not only the ones your model got wrong.
Match each review policy to its impact on validation-set fairness.
Put the label-review workflow in the right order to avoid bias.
What happens to the measured development accuracy if labels are corrected only for examples the model got wrong?
Checking only the dev examples your model got wrong is enough to guarantee an unbiased evaluation set.
Selective relabeling misses examples the classifier gets _____, so some bad labels are never examined.
Match each label-audit term with its meaning.
Put the steps in order to show why correcting only mislabeled mistakes can distort validation results.
Why does correcting only the labels of examples the model got wrong distort measured accuracy?
Why reviewing only mistaken labels can distort validation results
State the main risk of fixing only misclassified dev set labels.