Case Study

Check a Synthesized Audio Data Plan for Hidden Artifacts

Case context: A company is training a system to detect a particular machine fault from sound. To enlarge the dataset, the team records 30 minutes of factory ambient noise from one workshop and mixes that same noise into thousands of clean clips from different machines.

Question: Using the idea of representative synthesized training examples, what weakness should the team look for in this plan, and how should they improve it?

Sample answer: The weakness is that the synthesized set may be too easy for a model to identify as artificial because every mixed example is built from the same narrow noise source. That means the system may learn the repeated noise signature instead of the true fault-related audio patterns. The team should collect background noise from many workshops, machines, shifts, and acoustic conditions so the synthetic examples cover the variety seen in real use.

Key points:

  • A single background-noise recording can make all synthetic samples share the same artifact.
  • The model may learn to spot the repeated synthesis pattern rather than the target fault.
  • Representative synthesis requires source material that reflects real-world variation.
  • Broader noise collection is the appropriate fix.

Rubric: Full credit: identifies the single noise source as too narrow, explains the risk of learning a synthesis artifact, and recommends broadening the noise recordings. Partial credit: identifies only the narrow source or only the need for more diverse noise.

0

1

Updated 2026-08-12

Contributors are:

Who are from:

Tags

Machine Learning

Deep Learning

Supervised Learning

Dive into Deep Learning @ D2L

Data Science

Machine Learning Strategy

Machine Learning Yearning @ DeepLearning.AI