Using component error counts to decide where to improve a document-processing pipeline.
Case context: A team is building a document-processing pipeline with two stages: a page-segmentation module and a topic classifier. After reviewing 120 mistakes on the development set, they find that 84 errors come from the page-segmentation module and 36 errors come from the topic classifier.
Question: Which component should the team prioritize for improvement, and why?
Sample answer: The team should prioritize the page-segmentation module. It is responsible for most of the development-set mistakes, with 84 of the 120 errors, while the topic classifier accounts for only 36. Improving the component that causes the larger share of errors is more likely to reduce overall pipeline failures.
Key points:
- Prioritize the page-segmentation module.
- It causes 84 of the 120 observed errors.
- The topic classifier is responsible for the remaining 36 errors.
Rubric: The answer must recommend focusing on the page-segmentation module and justify that choice by noting that it is responsible for 84 of the 120 errors, whereas the topic classifier is responsible for 36.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Example Attribution Can Support a Second Pass of Debugging
Which component should receive the most attention if 72 of 80 dev-set mistakes come from the signal preprocessor?
A single source of most validation errors should receive priority.
Identify the component with the most assigned errors
Interpreting Component Error Counts
Order the steps for using component error counts to prioritize pipeline fixes.
A speech-recognition pipeline shows 9 times as many dev-set mistakes in the wake-word detector as in the speaker-id module. What should the team do next?
If one subsystem is responsible for 10 out of 100 validation mistakes, it should be the first area the team improves.
Reviewing 100 incorrect validation predictions and assigning each one to a pipeline _____ shows which step should be improved first.
Match each pipeline error-analysis term to its correct description.
Order the steps used to decide which pipeline component deserves improvement first.
Use component error counts to set improvement priorities.
Using component error counts to decide where to improve a document-processing pipeline.
Choose the component to improve when most dev-set mistakes come from one stage.