Bias from Correcting Only the Mistakes
If you change labels only for development examples that your system predicted incorrectly, the evaluation can become skewed rather than representative of the full development set.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Bias from Correcting Only the Mistakes
Which examples should be reviewed to improve dev-set label quality?
A single development example can have both an incorrect target label and an incorrect model prediction.
Review Both Error Cases and Correct Cases
Match each label-audit situation to the correct description.
Order the steps for a dev-set label audit.
Why can a dev example appear to be labeled correctly even when the label is wrong?
To assess label quality, it is enough to inspect only the examples your model got wrong.
Reviewing Labels on Development Examples
Match each label-review category to its role in checking data quality.
Why Correct Predictions Still Need Label Review
Why checking only mistakes can miss label problems
Why Reviewing Only the Flagged Errors Can Miss Label Problems
Why Recheck Apparently Correct Labels
Learn After
Reviewing Only Mistakes Can Skew a Dev Set
When Selective Relabeling Is Acceptable
What problem can occur if you correct only the dev-set examples your model got wrong?
True or False: If you relabel only the validation examples your system missed, the validation result remains an unbiased estimate.
Fixing labels only on the development examples your model misclassifies can introduce bias into evaluation because those examples are not selected at random.
How can selective relabeling skew a development-set score?
Editing Only the Wrongly Labeled Dev Examples Can Skew Your Evaluation
To avoid label-fix bias, review labels on _____ dev examples, not only the ones your model got wrong.
Match each review policy to its impact on validation-set fairness.
Put the label-review workflow in the right order to avoid bias.
What happens to the measured development accuracy if labels are corrected only for examples the model got wrong?
Checking only the dev examples your model got wrong is enough to guarantee an unbiased evaluation set.
Selective relabeling misses examples the classifier gets _____, so some bad labels are never examined.
Match each label-audit term with its meaning.
Put the steps in order to show why correcting only mislabeled mistakes can distort validation results.
Why does correcting only the labels of examples the model got wrong distort measured accuracy?
Why reviewing only mistaken labels can distort validation results
State the main risk of fixing only misclassified dev set labels.