Check a split strategy for a plant-disease app trained on mixed-source photos.
Case context: You are developing a phone app that identifies crop diseases. Your dataset has 90,000 lab-collected reference photos and 10,000 photos taken by farmers in real fields. A teammate proposes combining all 100,000 images and then drawing train, dev, and test splits at random so each split has the same overall mix.
Question: What is wrong with this splitting plan? Describe what the dev/test sets will mostly contain, and explain how that choice changes what the team ends up optimizing for.
Sample answer: The problem is that a random split of the combined pool will leave the dev and test sets dominated by the lab-reference images, at roughly 90% of each split. That does not match the real operating setting, where the app must perform well on farmer photos from field conditions. As a result, model tuning will be guided by performance on the easier lab distribution instead of the target field distribution.
Key points:
- Random splitting makes the dev/test sets about 90% lab photos.
- Lab-reference photos are not the same as the real target distribution.
- The team will tune the model for the wrong data source instead of the field images the app must handle.
Rubric: The response must state that the dev/test sets do not represent the target field-photo distribution, identify that approximately 90% of dev/test images will come from the lab source, and explain that model improvement will be directed toward the wrong distribution.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Why should source-mixed examples be kept out of both the development and test sets?
True or False: If your dataset mixes records from several sources with noticeably different distributions, the best practice is to randomly shuffle everything before creating the train, dev, and test splits.
For a movie-recommendation system, the dev and test sets should match the _____ of the users' future ratings, not a randomly shuffled archive.
If 188,000 website logs and 12,000 mobile-app logs are combined and then randomly split into dev/test sets, what fraction of the dev/test data will come from website logs?
If a model will be judged on roadside camera footage from one region, it is best to randomly mix together examples from several different regions when creating the dev and test sets.
Source-mix effect in a random split
Match each dev/test-set design choice to the effect it has on model development.
Put the reasoning steps in order for deciding whether mixed-source data should be randomly split into dev/test sets.
Selecting Validation and Test Data
If a dataset is formed by randomly mixing 4,000 photos from a mobile app with 156,000 photos collected from the web, the validation and test sets will reflect the mobile-app distribution.
Choose dev and test sets that match the data you expect to see in the _____.
Match each data-partitioning situation to the recommended outcome.
Order the steps for splitting mixed-source data so evaluation reflects the target distribution.
Analyze the consequences of mixing greenhouse and orchard photos before creating the dev and test sets.
Check a split strategy for a plant-disease app trained on mixed-source photos.
Why is a mixed dev/test split misleading in a warehouse-vision project?