Short Answer

Three Reasons a Held-Out Test Set Can Look Worse

Question: A model performs well on a tuning/validation set but much worse on a final test set that comes from another source. What three broad explanations should you consider?

Sample answer: The three main explanations are that the model was tuned too closely to the validation set, the final test set is genuinely more difficult, or the final test set comes from a different distribution than the validation set.

Key points:

  • Overfitting to the validation set
  • Final test set is harder
  • Final test set is different

Rubric: The answer must name all three explanations: overfitting to the validation set, a harder test set, and a different test set.

0

1

Updated 2026-08-12

Contributors are:

Who are from:

Tags

Machine Learning

Deep Learning

Machine Learning Strategy

Supervised Learning

Dive into Deep Learning @ D2L

Data Science

Machine Learning Yearning @ DeepLearning.AI