Case Study

How should a manual-audit sample be sized when records are plentiful?

Case context: A product analytics group can access a very large pool of records, but only a small number of staff are available to inspect mistakes by hand. The group is thinking about making its hand-audited development set much larger than its reviewers can realistically check.

Question: Identify the sizing error and name the main constraint that should determine the set size.

Sample answer: The error is choosing the hand-audited set size based on how much data exists instead of how much manual inspection the team can actually perform. The better rule is to size the set around the amount of time available for human review. In practice, teams rarely examine more than about 1,000 errors manually, so a very large set is not automatically better.

Key points:

  • A large data supply does not eliminate the limit on human review time.
  • The deciding factor should be the team’s realistic inspection capacity.
  • Making the set larger than reviewers can examine defeats the purpose of the guideline.
  • Manual examination of more than 1,000 errors is uncommon.

Rubric: The response should note that the set was sized too much by data availability, recommend using available manual-review time as the sizing constraint, and interpret the 1,000-error observation correctly.

0

1

Updated 2026-08-12

Contributors are:

Who are from:

Tags

Machine Learning

Deep Learning

Supervised Learning

Dive into Deep Learning @ D2L

Data Science

Machine Learning Strategy

Machine Learning Yearning @ DeepLearning.AI