Short Answer

What problem can occur if test-set scores are used to decide when to roll back a model version?

Question: If a team checks test-set performance every time it considers rolling back to an earlier model version, what is the main danger for the test set's value as an evaluation tool?

Sample answer: The main danger is that repeated use of the test set in decision-making will bias the team toward the test set. Over time, the test set stops being an independent check, so it can no longer provide a fully unbiased estimate of how the system will perform.

Key points:

  • Repeated decisions based on the test set bias the team toward that set
  • The test set loses its role as an independent evaluation
  • Its performance estimate is no longer fully unbiased

Rubric: The answer should explain that using the test set to guide rollback decisions risks biasing the team toward the test set, which prevents it from remaining a fully unbiased estimate of performance.

0

1

Updated 2026-08-12

Contributors are:

Who are from:

Tags

Machine Learning

Deep Learning

Machine Learning Strategy

Supervised Learning

Dive into Deep Learning @ D2L

Data Science

Machine Learning Yearning @ DeepLearning.AI