Estimating Human Error on Mobile Cat Images
To estimate human-level error on mobile cat images, ask people to label the same dataset and then measure how often they make mistakes.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Use the Strongest Practical Human Benchmark as a Proxy
Estimating Human Error on Mobile Cat Images
A Large Training-to-Validation Gap Signals High Variance
Why is the optimal error rate for detecting a bicycle in a clear photo close to 0%?
If 8% of satellite images are impossible for experts to label correctly because thick cloud cover hides the target, the best achievable classification error is approximately 8%.
Human-level performance is often used as a proxy for the _____ error rate on a task.
Match each situation with the approximate best possible error rate suggested by human performance.
Order the steps for using human performance as a stand-in for the best achievable error rate and deciding whether to reduce bias.
A defect detector has 12% error, while expert annotators have 4% error. What is the avoidable bias and what should you do next?
Assuming a Zero Error Floor
In a task where people almost never miss the correct answer, the ideal error rate is nearly _____.
Match each term to its meaning when human performance is used as a practical benchmark for model error.
Order the steps for deciding whether a task has a near-zero or a much higher best-possible error rate.
How model and expert error comparison reveals bias
Selecting a Desired Error Rate Proxy for a Noisy Dictation System
What information should a person use when checking one stage of a pipeline against human performance?
Learn After
Measuring Human Error on Phone Photos
Estimating Human Performance on a Domain-Specific Image Task
Estimating Human Error on an Image Set
Parts of an Evaluation on the Target Dataset
Order the steps for estimating human accuracy on phone photos.
How to Measure Human Performance on the Target Data Set
Choosing a Benchmark for Night-Time Phone Photos
Human Labeling for Performance Checks
Benchmarking Human Performance
Human Error Measured on the Same Image Type