Mismatched Auxiliary Data Source
An auxiliary data source is inconsistent with the main task when the same feature values can point to different labels depending on which source the example came from. For example, if the goal is to predict home prices in Denver, a dataset from Phoenix may be inconsistent because the same square footage and bedroom count can correspond to a different price range in that market. Mixing such data without adjustment can confuse the model and reduce performance on the target task.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Some Data Should Be Left Out of Training
When Development and Test Sets Reflect Different Populations
How Model Capacity Changes the Risk of Mixing Data Sources
One Predictor Can Work Across Multiple Data Sources
Choose evaluation data to match the real-world target
Mismatched Auxiliary Data Source
Building Dev and Test Sets Before Real Users Exist
Refreshing Evaluation Sets After a Product Launch
Using Public Web Images When No Better Future-Like Data Exists
Judging How Much to Invest in Dev and Test Sets
What should determine dev and test set selection?
True or False: You can assume the training set and test set always come from the same distribution.
Development and test sets should reflect the conditions you expect after deployment, not only the _____ available in your training pool.
Why can a simple random test split be a poor choice when the data you expect in the future is different from the data you have now?
You can usually assume the data used for training and the data used for testing come from the same distribution.
Design Dev and Test Sets for the Future
Match each concept about development and test sets to its description.
Order the steps for choosing development and test sets when future data differs from training data.
What should dev and test examples be designed to resemble?
A validation and test set must exactly match the training distribution in every project.
How should a test set be chosen when deployment data will differ?
Match each data scenario to the best dev/test set choice.
Order the reasoning steps for deciding whether a dev/test split is appropriate.
Why a Random 30 Percent Split Can Be Misleading When Future Data Will Differ
Dev and Test Splits for a Field-Photo Classifier
How to Choose Dev and Test Data When Future Data Will Differ
Learn After
Adding a Source-ID Feature for Conflicting Data
Training a rent predictor with records from two neighborhoods that use different pricing rules
Rent data for studio apartments from two different cities can be treated as consistent just because the apartments have the same floor area.
When task data conflicts with the target domain, _____ the mismatched examples during training.
Match the terms in a data-shift example
How to decide whether to include data from a second source
When Does Auxiliary Data Conflict with the Target Task?
Combining datasets with conflicting labels can hurt model performance
Relative pricing of a suburban home compared with _____ homes
Match each example to the correct consistency category
Order the steps that explain why combining two sources with different label rules can hurt learning.
When Auxiliary Data Conflicts with the Target Task
Deciding whether to add auxiliary rent data from another city
What makes an auxiliary data source inconsistent with the main task?