How should a team use strong human performance to guide an ML system with substantial error?
Case context: A team is developing an ML system for a task that people perform well. Human labelers can produce examples, and reviewers can understand many of the system’s mistakes. The algorithm’s error remains well above the human reference.
Question: Diagnose what the comparison suggests and describe how the team should use human performance during development.
Sample answer: The gap between the algorithm and the human reference suggests high avoidable bias. The team should use human-level performance to estimate optimal error and set a reasonable, achievable desired error rate. It can then draw on human labelers for data and human intuition for error analysis. Because avoidable bias is high, the team has a menu of improvement options to explore.
Key points:
- The performance gap indicates high avoidable bias.
- Human-level performance informs optimal and desired error rates.
- The target should be reasonable and achievable.
- Human labelers can provide data.
- Human intuition can guide error analysis.
- High avoidable bias opens improvement options.
Rubric: The response should diagnose high avoidable bias, use the human reference to define optimal and desired error rates, and explain the roles of human labeling and intuition. It should connect the diagnosis to available improvement options.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Human Labelers Make Data Easier for Human-Solvable Tasks
Human Intuition Can Guide Error Analysis
Human-Better Data Subsets Can Drive Progress After Surpassing Average Human Performance
Which combination explains why human-level comparison can make ML development easier?
Human-level performance can help estimate optimal error and establish a desired error rate.
A reasonable and achievable target error rate can accelerate a team’s _____.
Match each human-level comparison benefit with its role in ML development.
Order the reasoning process for using human-level performance to guide development.
Explain how human-level comparison supports both diagnosis and team progress in ML development.
How should a team use strong human performance to guide an ML system with substantial error?
Why is identifying high avoidable bias valuable to an ML team?
Which target would best support faster team progress according to the source?
Human-level comparison is useful only for obtaining labeled data.