Explain why combining two existing audio sources can meet an in-car speech-data need.
Question: In a concise response, explain how quiet-room speech and car or road noise can be used to create the desired training audio and why this approach may be useful.
Sample answer: A speech system may need more data that sounds as if it were recorded inside a car. Instead of collecting a large amount of speech while driving around, one can obtain car or road noise clips and add them to recordings of people speaking in a quiet room. The combined clips sound like people speaking in noisy cars, making artificial synthesis an easier way to produce the desired kind of data.
Key points:
- The target is speech that sounds recorded inside a car.
- The inputs are quiet-room speech and car or road noise.
- The noise is added to the speech audio.
- The result sounds like a person speaking in a noisy car.
- Synthesis may be easier than extensive collection while driving.
Rubric: A strong response identifies the desired in-car sound, names both required audio sources, explains that they are added together, describes the resulting noisy-car speech, and contrasts synthesis with collecting substantial data while driving.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Which synthesis method can create speech that sounds recorded inside a noisy car?
Can artificial synthesis reduce the need to collect large amounts of speech while driving?
Adding quiet-room speech to _____ can produce noisy-car speech audio.
Match each element of noisy-car speech synthesis to its role.
Order the reasoning process for synthesizing noisy-car speech data.
Explain why combining two existing audio sources can meet an in-car speech-data need.
Decide how a speech team should create more audio that sounds recorded in a car.
What two audio ingredients are needed for the described noisy-car speech synthesis?
A team has quiet speech and road-noise clips. What output should combining them produce?
Does road-noise audio alone satisfy the need for in-car speech examples?