How do unusually good and bad small subsets affect learning-curve points?
Question: Answer in one to three sentences.
Sample answer: A particularly bad small subset, such as one with many ambiguous or mislabeled examples, can produce a worse-than-expected learning-curve value. A particularly good subset can produce a better-than-expected value, so the curve fluctuates at small training-set sizes.
Key points:
- Bad subsets can yield worse-than-expected values
- Good subsets can yield better-than-expected values
- Random variation is more visible at small training-set sizes
Rubric: The answer should contrast the effects of unusually good and bad subsets and connect both effects to fluctuation in learning-curve values.
0
1
References
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Averaging Learning Curves over Multiple Random Subsets
Balanced Subsets for Noisy Learning Curves in Skewed or Many-Class Data
Why can a learning-curve point fluctuate when it is based on a very small random training subset?
A small random subset can produce a learning-curve value that is higher or lower than expected.
A small subset with many ambiguous or mislabeled examples is unusually _____.
Match each small-subset condition to its learning-curve implication.
Order the reasoning used to diagnose a noisy point at a small training-set size.
Explain why small training subsets can make learning curves noisy.
Diagnose a noisy learning-curve point from a ten-example subset.
How do unusually good and bad small subsets affect learning-curve points?
Which dataset condition most increases the risk that a tiny random subset will be unrepresentative?
With many classes, a small random subset is less likely to be unrepresentative.