Analyze the consequences of mixing greenhouse and orchard photos before creating the dev and test sets.
Question: Suppose a plant-disease team has 155,000 labeled images: 150,000 from a greenhouse camera system and 5,000 from farmers' phone uploads. If all 155,000 images are randomly shuffled before splitting into train, dev, and test sets, what problems would this create for evaluation? Why does that split strategy conflict with the goal of choosing dev and test data?
Sample answer: Random shuffling would make the dev and test sets come from the same overall mixture as the training set, but that mixture is dominated by greenhouse images. Since 150,000 of 155,000 images are greenhouse photos, about 96.8% of the dev and test examples would also be greenhouse images. That means the evaluation sets would mainly measure performance on greenhouse conditions, even if the real deployment target is farmers' phone photos. The team could end up optimizing decisions for the easier or more common source instead of the data it actually needs to handle in production.
Key points:
- Random shuffling forces train, dev, and test to reflect the same blended distribution.
- With 150,000 of 155,000 images coming from the greenhouse source, dev and test would be about 96.8% greenhouse images.
- The resulting evaluation would not match the target distribution of farmer-uploaded photos.
- This breaks the rule that dev and test sets should reflect the data the system must perform well on in the future.
Rubric: The answer must explain that random shuffling makes dev and test sets dominated by greenhouse images (about 96.8%), that this does not match the intended phone-photo distribution, and that this mismatch violates the principle that evaluation data should resemble future deployment data.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Why should source-mixed examples be kept out of both the development and test sets?
True or False: If your dataset mixes records from several sources with noticeably different distributions, the best practice is to randomly shuffle everything before creating the train, dev, and test splits.
For a movie-recommendation system, the dev and test sets should match the _____ of the users' future ratings, not a randomly shuffled archive.
If 188,000 website logs and 12,000 mobile-app logs are combined and then randomly split into dev/test sets, what fraction of the dev/test data will come from website logs?
If a model will be judged on roadside camera footage from one region, it is best to randomly mix together examples from several different regions when creating the dev and test sets.
Source-mix effect in a random split
Match each dev/test-set design choice to the effect it has on model development.
Put the reasoning steps in order for deciding whether mixed-source data should be randomly split into dev/test sets.
Selecting Validation and Test Data
If a dataset is formed by randomly mixing 4,000 photos from a mobile app with 156,000 photos collected from the web, the validation and test sets will reflect the mobile-app distribution.
Choose dev and test sets that match the data you expect to see in the _____.
Match each data-partitioning situation to the recommended outcome.
Order the steps for splitting mixed-source data so evaluation reflects the target distribution.
Analyze the consequences of mixing greenhouse and orchard photos before creating the dev and test sets.
Check a split strategy for a plant-disease app trained on mixed-source photos.
Why is a mixed dev/test split misleading in a warehouse-vision project?