Averaging Learning Curves over Multiple Random Subsets
If learning-curve noise makes the true trend hard to see, one remedy is to sample several different training sets of the same small size, train a different model on each, compute each model's training and dev error, and plot the average training and dev error.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Averaging Learning Curves over Multiple Random Subsets
Balanced Subsets for Noisy Learning Curves in Skewed or Many-Class Data
Why can a learning-curve point fluctuate when it is based on a very small random training subset?
A small random subset can produce a learning-curve value that is higher or lower than expected.
A small subset with many ambiguous or mislabeled examples is unusually _____.
Match each small-subset condition to its learning-curve implication.
Order the reasoning used to diagnose a noisy point at a small training-set size.
Explain why small training subsets can make learning curves noisy.
Diagnose a noisy learning-curve point from a ten-example subset.
How do unusually good and bad small subsets affect learning-curve points?
Which dataset condition most increases the risk that a tiny random subset will be unrepresentative?
With many classes, a small random subset is less likely to be unrepresentative.
Learn After
When to Use Noise-Reduction Techniques for Learning Curves
When averaging learning curves over multiple random subsets, what is the correct procedure for each randomly selected small training set?
To reduce noise in a learning curve, you should train a single model on multiple randomly chosen training sets of the same small size.
After training a different model on each random subset of the same small size, you compute and plot the _____ training error and dev set error.
Why does averaging learning curves over multiple random subsets help reveal the true learning trend?
The averaging technique described by Andrew Ng uses sampling with replacement to create the multiple small training sets.
Instead of training just one model on a small set, Ng recommends training _____ different models on different randomly chosen subsets of the same size.
Match each term in the averaging technique to its correct description.
Order the steps of the averaging-over-multiple-random-subsets procedure as described in Machine Learning Yearning.
According to Machine Learning Yearning, approximately how many randomly chosen training subsets should you select when using the averaging technique?
In the averaging technique, a single model is trained jointly on all the randomly chosen small training subsets combined into one larger set.
After training each model on a small subset, you compute both the training error and the _____ error for each model before averaging.
Match each challenge to the element of the averaging technique that directly addresses it.
Order the reasoning steps a practitioner should follow when deciding to apply and interpret the averaging technique for noisy learning curves.
How does averaging learning curves over multiple random subsets reduce noise, and what are the specific steps to execute this technique?
Smoothing a Noisy Dev Set Learning Curve at Small Training Sizes
Sampling Method and Metric Plotting for Learning Curve Noise Reduction