Case Study

Should a new product team spend three weeks polishing its first evaluation set?

Case context: A startup has been asked to build a fraud-detection model for a new mobile-wallet product. There are no previous benchmarks or evaluation sets, and one group proposes spending three weeks designing the first dev/test split and scoring metric so the setup is as polished as possible before any modeling begins.

Question: Should the team wait three weeks before starting model development? What is the better approach, and why?

Sample answer: No. For a brand-new project, the first dev/test set and metric should be assembled quickly, ideally in less than a week. That gives the team a concrete target and lets them begin learning from experiments right away. Spending several weeks perfecting the setup slows iteration more than it helps. Once the system has real data and the task is better understood, the team can revisit the evaluation design and invest much more time in improving it.

Key points:

  • Three weeks is longer than recommended for an early-stage project
  • A rough but fast initial evaluation setup is enough to start
  • Early setup should create a clear target for experimentation
  • The evaluation design can be improved later when the project is more mature
  • Excessive early planning delays useful model iteration

Rubric: Full credit says not to wait three weeks, states that the initial dev/test set and metric should be built in under a week, and explains that the setup can be refined later after the project matures. Partial credit mentions the delay problem without clearly tying it to the early-stage guideline.

0

1

Updated 2026-08-12

Contributors are:

Who are from:

Tags

Machine Learning

Deep Learning

Supervised Learning

Dive into Deep Learning @ D2L

Data Science

Machine Learning Strategy

Machine Learning Yearning @ DeepLearning.AI