Why Human Baselines Are Less Helpful on Hard-to-Judge Tasks
When a task is one that people do poorly, such as ranking restaurant recommendations, forecasting weekly demand, or estimating loan default risk, a human-level benchmark is much less informative. In those settings, human labels may be unreliable, intuition is a weak guide, and it is hard to define a meaningful target error rate.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Why Human Baselines Are Less Helpful on Hard-to-Judge Tasks
Which situation makes it hard to estimate the best achievable error rate for a task?
Recommending products to shoppers is an example of a task that people may find hard to evaluate well.
When human performance is poor, estimating the _____ error rate becomes difficult.
Match each concept to the description that best fits how its error rate can be estimated.
Order the steps for judging whether expert performance can set a target error rate.
Which task is given as an example of a problem that is hard for people to judge reliably?
Human performance is always a dependable estimate of the best achievable error rate, even when people find the task difficult.
Predicting Which _____ to Show a Shopper Can Be Hard for People Too
Match Each Situation to Its Effect on Estimating the Best Possible Error
Order the reasoning steps for a task that is difficult even for people
Why Human Performance May Not Give a Useful Error Target
Estimating a Benchmark for Ticket Routing
Examples where human judgment is a poor guide to optimal error
Learn After
Why is a human benchmark often weak for movie recommendations, credit-risk scoring, or demand forecasting?
True or False: It is difficult for people to assign the single best movie recommendation to each user in a database.
A classic hard forecasting example is the _____.
Match each situation where human performance is weak to the reason it matters for machine learning.
Order the reasoning steps for why a weak human benchmark makes evaluation difficult.
Why a weak human baseline is a poor yardstick in some prediction tasks
Explain why a podcast recommender has no obvious error-rate goal.
Give Two Examples of Tasks People Handle Poorly
Why is human intuition a weak guide for building a stock-prediction model?
True or False: A human-level benchmark is equally informative for every machine learning problem.