Learn Before
Why Human Comparison Becomes Less Helpful After Strong Model Performance
When people can no longer easily tell which examples the model is still getting wrong, only some human-comparison methods remain useful. At that point, progress usually becomes slower on tasks where the model already exceeds human performance, because humans have fewer reliable ways to spot weaknesses. When the model is still behind people, those comparison methods tend to be more informative and improvement is often faster.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Why Human Comparison Becomes Less Helpful After Strong Model Performance
When can human comparison still be useful after a model is already better than average human performance on the dev or test set?
Human judgment can still be useful on specific subsets even after a model beats average human performance overall.
When people still do better than the model on certain cases, they can provide better _____, useful intuition, and a target level of performance.
Match each benefit of comparing a model against people on a hard subset to its description.
Deciding What to Do After Beating Average Human Performance
In the document-reading example, on which task does the system beat people while people still do better on a different task?
If a classifier beats the average human score on the development set, then comparing its mistakes with human judgments is no longer useful for finding improvements.
Human-comparison methods remain useful as long as there are dev examples where people are _____ and the model is mistaken.
Match each item from the invoice-reading example to its role in the human-better-subset idea.
Order the steps for using a human-strong subset to improve a model.
How human comparison can still help a stronger-than-average system
Using a human-strong subset to improve a transaction review model
When Human-Versus-Model Comparisons Stop Helping
Learn After
When does human comparison stop helping much?
Faster Gains Before Surpassing a Human Baseline
When people can no longer easily point out the model’s obvious mistakes, only a _____ of comparison methods remain useful.
Match each situation to the usual direction of progress.
Put the reasoning steps in order for why improvement slows after machine performance exceeds human performance.
Why improvement slows once a model beats human performance
Explain why progress has become harder for a package-screening model.
Examples of Tasks Where Machines Already Excel
When are human-checking methods still especially useful during model development?
True or False: Once a model clearly outperforms people on a task, every human-based evaluation method remains equally useful.