Learn Before
Splitting a large development set into a review subset and a tuning subset
When a development set is large, manual inspection can become a bottleneck. For example, if a set contains 8,000 cases and the model makes mistakes on about 15% of them, there may be roughly 1,200 errors to inspect. Rather than examining all of them by hand, the team can divide the development set into two parts. One part is reserved for human review of errors, and the other part is kept untouched for choosing model settings. The reviewed portion will be adapted to more quickly, so separating the two roles makes it easier to notice when manual analysis is starting to bias decisions. The untouched portion then provides a cleaner signal for tuning.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
One Example May Fit Several Error Tags
New Error Categories Can Appear During Review
Choose Error Categories You Can Act On
Error Review Improves Through Repeated Passes
Using Error Counts to Decide Where to Focus Next
Working on Several Error Buckets at Once
Error Analysis Is Not an Automatic Ranking Rule
A Category's Share of Errors Sets an Upper Bound on Improvement
Error Analysis Helps Estimate Whether a Proposed Change Is Worth the Effort
Why Quick Error Review Is Often Skipped
Incorrect Labels in a Validation Set
Splitting a large development set into a review subset and a tuning subset
Build a Simple Baseline First, Then Use Error Analysis to Prioritize Improvements
Using Training-Set Mistakes to Diagnose High Bias
Reviewing a Sample of Validation Errors
Separating Search Errors from Scoring Errors in Inference
Component-Wise Error Review
Error Analysis as a Data-Science Lens on Model Mistakes
Multiple Valid Approaches to Error Analysis
Tasks Humans Can Perform Give Stronger Error Analysis Benchmarks
When diagnosing a model, what should error analysis focus on first?
Error analysis on a machine learning system must follow one fixed procedure.
Name the practice of reviewing mistakes to understand why predictions failed.
Match each error-analysis idea to the description that best fits it.
Put the steps of a simple dev-set error review in the right order.
Why is it useful to inspect misclassified examples during error analysis, even for error types you cannot immediately repair?
Error analysis is usually repeated after each round of model changes.
Error analysis can help you judge which improvement paths look most _____.
Match each error-analysis activity with the benefit it can provide.
Order the steps for deciding which error types to target after a first pass of error review.
Why Error Review Helps Set the Right Next Priorities
Plan the next review step after repeated image-classifier mistakes.
What is error analysis used for in machine learning?
Learn After
Human-Review Dev Set
Blackbox Dev Set
Use the Full Dev Set When It Is Too Small to Split
Why split a development set into a review subset and a tuning subset?
Does examining part of a dev set more closely increase the risk of overfitting to it?
The hands-off portion of the dev set can still be used to tune _____.
Match each dev-set idea with its role or consequence.
Order the reasoning for managing a dev set that is too large to inspect fully by hand.
Explain how separating a reviewed subset from an untouched subset can reveal overfitting during error analysis.
How should a team organize a large validation set with many mistakes?
How are the two parts of a split dev set used?
Which result suggests a model has been tuned too closely to the hand-checked subset?
If a development set is too small to divide into separate analysis and tuning subsets, using the full set for both purposes is a reasonable choice.