Essay

How should the choice of human reference change as a model gets better?

Question: A team is evaluating a classifier with a current error rate and must decide what kind of human benchmark to use for future labeling and model improvement. Explain how that choice should differ when the system is still far from acceptable performance versus when it is already close to human performance. Use an example with two reviewers who have noticeably different error rates.

Sample answer: When the model is still making many mistakes, the exact human benchmark is less critical because the system is so far from good performance that either reference gives similar guidance. For instance, if a quality-check model is at 36% error, comparing it with a reviewer at 11% error or one at 4% error will not change the main conclusion that there is a lot of room to improve. Once the model gets much stronger, though, a rough benchmark is no longer enough. If the model has improved to about 9% error, then a reviewer at 11% error would be too weak a reference, while a reviewer at 4% error gives a much better target and more useful feedback for the next round of improvements.

Key points:

  • When system error is large, the precision of the human reference matters less.
  • When system error is small, the reference should be much more accurate.
  • A better human reference is needed to guide fine-grained improvement near strong performance.

Rubric: Full credit is given for explaining that a model with high error can be judged with a less precise human benchmark, while a model with low error needs a much more accurate benchmark to provide useful guidance, supported by a different but comparable example.

0

1

Updated 2026-08-12

Contributors are:

Who are from:

Tags

Machine Learning

Deep Learning

Supervised Learning

Dive into Deep Learning @ D2L

Data Science

Machine Learning Strategy

Machine Learning Yearning @ DeepLearning.AI