Synthetic Data That Approximates the Dev Distribution
Synthetic data generation can help you build a very large dataset that reasonably matches the dev set distribution.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Data Mismatch May Not Have a Clear Fix
Synthetic Data That Approximates the Dev Distribution
Are Synthetic Training Examples Representative?
What should you do when training results are strong but validation results drop because the data sources differ?
A data mismatch problem happens when a model does well on training data but performs poorly on a dev set that comes from a different distribution.
Matching the Dev Set Environment
Match each data-mismatch idea to its description.
Order the steps for diagnosing and fixing a data mismatch in a customer-feedback classifier.
What most likely explains the model’s weak performance on the development set in this speech project?
Does training on examples that look more like the dev set always eliminate data mismatch?
When a speech recognizer performs poorly on noisy clips in the dev set, one remedy is to collect more training data that better _____ those difficult examples.
Match each part of a traffic-sign recognition scenario to its role in a data mismatch diagnosis.
Order the steps for deciding whether to collect training data that better matches difficult dev examples.
Using Targeted Data to Reduce Distribution Mismatch
Fix a training-dev distribution gap
What training data change helps with a data mismatch problem?
Learn After
Creating Subway-Like Speech by Mixing Clean Voice and Transit Noise
Adding Artificial Blur to Training Photos
A Gap Between Human and Machine Judgments of Synthetic Data
When Synthetic Data Becomes Useful
What does artificial data synthesis help you build when your development set is missing important cases?
True or False: If your validation data contains an important rare pattern, generating synthetic examples can help enlarge the training set so it better covers that pattern.
Artificial data synthesis can help create a _____ that better matches the validation set.
What is the main advantage of synthetic data when your training set does not match the dev set?
Artificially generated examples always match the dev set’s real-world distribution exactly.
Artificially generated examples can help create a _____ dataset that still resembles the development set.
Match each synthetic-data situation to the real-world factor it is meant to imitate.
Order the reasoning steps for deciding whether generated data can help match a validation distribution.
When is synthetic data most useful for matching a development set?
Synthetic examples can help narrow the difference between training data and development data distributions.
There are several _____ in which artificial data generation can produce a large dataset that closely matches the development set.
Match each concept in synthetic-data design to its best description.
Order the steps for creating synthetic office-call audio to resemble a noisy support-center dev set.
When is synthetic data useful for matching a development set?
When Synthetic Data Is Worth Building for a Narrow Validation Set
What should synthesized training data achieve when it is built to mirror a dev set?