Essay

Using a Human Benchmark to Choose Where to Improve a Model

Question: Why is it useful to compare each stage of a machine learning system with expert human performance before choosing what to improve next? Explain why a stage that is far below the human benchmark often deserves priority.

Sample answer: A human benchmark gives a practical reference point for judging how much room a system still has to improve. When one stage is much worse than expert human performance, that gap suggests there is still substantial headroom and that better results may be achievable with focused work. By contrast, if a stage is already close to the human benchmark, further gains are usually smaller and harder to obtain. For that reason, engineers often prioritize the component with the largest gap because effort spent there is more likely to produce meaningful improvement.

Key points:

  • Human performance provides a reference for estimating remaining improvement potential.
  • A large gap usually signals a promising target with more attainable gains.
  • Prioritizing the weakest stage helps avoid spending effort where returns are likely to be small.

Rubric: The response should explain 1) that human performance serves as a benchmark, 2) that a large gap indicates higher improvement potential and thus higher priority, and 3) that this approach helps avoid overinvesting in stages that are already near the benchmark.

0

1

Updated 2026-08-12

Contributors are:

Who are from:

Tags

Machine Learning

Deep Learning

Supervised Learning

Dive into Deep Learning @ D2L

Data Science

Machine Learning Strategy

Machine Learning Yearning @ DeepLearning.AI