Case Study

How should a team respond when its evaluation score no longer reflects real-world usefulness?

Case context: A speech-recognition team notices that its main evaluation score treats every mistake the same way. That becomes a problem because some mistakes are minor, while others cause serious confusion for users. The team no longer trusts the score as a measure of real progress. One group suggests picking the best-looking model by hand after reviewing examples, while another group argues for designing a better metric.

Question: Assess the two options and explain what the team should do next. Describe how the choice affects day-to-day work and how the team's objective should be defined.

Sample answer: The team should avoid relying on manual model-by-model selection as the long-term solution. Instead, it should create a new evaluation metric that matches the real goal more closely, such as one that penalizes the most harmful mistakes more heavily. That metric should then become the team's explicit target. This gives everyone an automatic objective to optimize and keeps progress aligned with the product goal.

Key points:

  • Rejects manual selection as the main strategy.
  • Recommends building a new metric that better matches the desired outcome.
  • Uses that metric to define a clear, automated team goal.

Rubric: A correct response must say that manual selection should not be the primary approach and should recommend creating a new metric that becomes the team's explicit goal.

0

1

Updated 2026-08-12

Contributors are:

Who are from:

Tags

Machine Learning

Deep Learning

Machine Learning Strategy

Supervised Learning

Dive into Deep Learning @ D2L

Data Science

Machine Learning Yearning @ DeepLearning.AI