Why a tighter human reference helps when error is already low
Question: A model has improved to about 8% error. Why would it still be useful to define a more careful human benchmark at 2% error instead of using a looser reference?
Sample answer: A tighter human benchmark gives a more informative target when the model is already fairly strong, so it helps the team see more clearly how much room for improvement remains and how to keep making progress.
Key points:
- A precise human-level reference matters more when system error is already low.
- It gives better guidance for the next round of improvements.
Rubric: The answer must explain that a more precise human benchmark provides better guidance or a better target for continuing to improve an already strong system.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
How Human Performance Guides Machine Learning Work
At 28% system error, which benchmark is most useful for deciding where to improve next?
If a classifier has about 40% error, using a nurse practitioner with 12% error instead of a senior specialist with 6% error as the human reference makes only a small practical difference.
Using a human benchmark when error is already fairly low
Match each model error situation to the lesson it gives about choosing a human benchmark.
Order the steps for deciding when a more precise human benchmark is worth using.
Why is a 2% human benchmark more useful for guiding improvement when a system has 10% error than when it has 40% error?
At a 40% system error rate, switching between a 12% human benchmark and a 6% human benchmark usually changes the diagnosis a great deal.
When a model is meant to match expert inspectors, their error rate can be the _____ error rate for the system.
Match each labeler or system case to its description.
Order the steps for choosing a human performance reference when evaluating an ML system.
How should the choice of human reference change as a model gets better?
Select the right human reference when system error is very high.
Why a tighter human reference helps when error is already low