Case Study

Selecting a Desired Error Rate Proxy for a Noisy Dictation System

Case context: A company is building a voice-to-text system for field inspectors. During review, they discover that 11% of the recorded clips are so distorted by wind and machinery that even trained human transcribers cannot reliably understand them. The current model has a 13% transcription error rate.

Question: Based on this case, what should the team use as its desired error rate proxy, and how should that shape its performance target for the model?

Sample answer: The team should use about 11% as the desired error rate proxy, because that portion of the data is effectively impossible to transcribe accurately under the current recording conditions. Since the model is already at 13%, the gap to human-level performance is only about 2%. That means the team should not expect a near-zero error rate from this dataset alone; instead, it should focus on reducing the remaining 2% gap or improving the quality of the audio.

Key points:

  • Use 11% as the proxy for the best achievable error rate on this data.
  • The avoidable gap is roughly 2%.
  • A near-zero error target is unrealistic without better input quality.

Rubric: A strong response will identify 11% as the proxy for the best achievable error rate, compute the gap as 2%, and conclude that aiming for 0% error on this dataset is not realistic.

0

1

Updated 2026-08-12

Contributors are:

Who are from:

Tags

Machine Learning

Deep Learning

Supervised Learning

Dive into Deep Learning @ D2L

Data Science

Machine Learning Strategy

Machine Learning Yearning @ DeepLearning.AI