Concept icon
Concept

The Best Human Benchmark Depends on How Accurate the System Already Is

When a model is still making a large number of mistakes, small differences between human references may not change the improvement plan much. For instance, if a classifier has 38% error, using a clinician with 11% error instead of one with 4% error may not affect the next debugging step very much. But if the model is already near 9% error, comparing against a 2% human reference gives a much clearer target for further gains.

0

1

Concept icon
Updated 2026-08-12

Contributors are:

Who are from:

Tags

Machine Learning

Deep Learning

Supervised Learning

Dive into Deep Learning @ D2L

Data Science

Machine Learning Strategy

Machine Learning Yearning @ DeepLearning.AI