Case Study

A team has abundant data but limited review time. How should it size its Eyeball dev set?

Case context: A machine learning team can draw from a plentiful supply of examples. However, its members have limited time to inspect errors manually and are considering creating an Eyeball dev set larger than they can review.

Question: Diagnose the team's sizing mistake and state the source-grounded factor it should use instead.

Sample answer: The mistake is sizing the Eyeball dev set according to the abundant supply of data rather than the team's capacity to inspect examples. The team should base the size mainly on how many examples it has time to analyze manually. It should also recognize that manually analyzing more than 1,000 errors is rare, rather than treating a very large set as automatically desirable.

Key points:

  • Abundant data does not remove the manual-review constraint.
  • The set should reflect realistic manual-analysis capacity.
  • A set larger than the team can inspect conflicts with the stated sizing principle.
  • Manual analysis of more than 1,000 errors is described as rare.

Rubric: The response should identify the mismatch between set size and review capacity, recommend sizing by available manual-analysis time, and interpret the 1,000-error observation accurately.

0

1

Updated 2026-07-19

Contributors are:

Who are from:

Tags

Machine Learning

Deep Learning

Supervised Learning

Dive into Deep Learning @ D2L

Data Science

Machine Learning Strategy

Machine Learning Yearning @ DeepLearning.AI