Learn Before
Spotting Overfitting to a Manually Reviewed Dev Slice
When people repeatedly inspect a subset of the dev set to guide error analysis, they can unconsciously tune decisions to that reviewed subset. If scores rise much faster on the reviewed subset than on a separate hidden subset used for comparison, that is a sign the reviewed subset has been overfit. Dividing the dev set into a manually examined portion and a separate untouched portion makes it possible to detect this problem and keep the analysis process honest.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Spotting Overfitting to a Manually Reviewed Dev Slice
How to Size an Eyeball Dev Set When Data Is Easy to Review
A Review Set Should Expose the Main Error Types
Manual Review Sets Can Be Unhelpful for Tasks People Cannot Judge Reliably
What is the main purpose of the human-review subset when a validation set is split into two parts?
The label "manual review set" can remind a team that people should inspect those examples directly.
What fraction of the dev set is manually reviewed in the Eyeball subset?
Match each dev set split term to its description.
Arrange the steps for building and using a small review subset from a larger development set in the correct order.
If 20% of a 3,000-example dev set is set aside as the Eyeball dev set, how many examples are in that subset?
If a 400-example Eyeball dev set has the same 20% error rate as the full dev set, you would expect about 80 misclassified examples.
The eye-check dev set should contain enough mistakes for you to _____.
Match each manual-review subset fact to the idea it describes.
Order the steps for deciding whether a small inspection subset is large enough for error analysis.
How should the manual-review subset be named in a voice transcription project?
Split the Development Set Into an Inspection Set and a Locked Set
How large should an Eyeball dev set be for useful manual error review?
Learn After
Overfitting a development set used for model tuning
What most clearly indicates that a manually inspected dev slice has been over-tuned during error analysis?
Inspecting a manually reviewed validation set can make you adapt to that set more quickly than if you never looked at its examples.
Keeping a Reviewed Subset Separate from a Hidden Check Set
Match each validation-set concept to its description in a reviewed-vs-reserved split.
Order the steps for checking whether repeated manual review is causing overfitting to a reviewed validation subset.
What is the most likely conclusion when manual review keeps pushing one dev set score far above another?
In the Eyeball/Blackbox approach, the Blackbox dev set is checked manually during routine error analysis.
If the Eyeball dev set improves much faster than the Blackbox dev set, you have _____ the Eyeball dev set.
Match each performance pattern with the most appropriate interpretation when comparing a hand-reviewed set with a separate hidden set.
Order the steps for checking whether a manually reviewed dev set has been overfit.
Detect Overfitting to an Inspected Development Set
When Manual Review Improves Faster Than a Hidden Evaluation Set
Why split the validation set into reviewed and unreviewed parts?