Case Study

Explain why a stronger dev result may not justify the next engineering priority.

Case context: A product team evaluates models using a dev set and a test set that come from different user populations. A recent model change improves the dev score, but the team is unsure whether similar work should be scheduled next.

Question: What issue should the team diagnose as the reason for its hesitation, and why does that make prioritization difficult?

Sample answer: The team should diagnose a mismatch between the dev and test distributions. An improvement on the dev set does not guarantee an improvement on the test set, because the two sets reflect different data. As a result, the team cannot be sure the change will help the final target population, which makes it hard to decide whether related work should be the next priority.

Key points:

  • The dev set and test set come from different distributions.
  • Better dev performance may not carry over to test performance.
  • The team is uncertain whether the change truly helps the target data.
  • That uncertainty weakens the basis for deciding what to prioritize.

Rubric: The response should identify the dev/test distribution mismatch, explain why dev gains may fail to transfer to test performance, and connect that uncertainty to difficulty choosing priorities.

0

1

Updated 2026-08-12

Contributors are:

Who are from:

Tags

Machine Learning

Deep Learning

Supervised Learning

Dive into Deep Learning @ D2L

Data Science

Machine Learning Strategy

Machine Learning Yearning @ DeepLearning.AI