Short Answer

Input and Output in a Speech-to-Text System

Question: In a speech-to-text system, what is the input and what is the output? Also, how does that output differ from a simpler prediction such as a single class label or a numeric score?

Sample answer: The input is the audio recording, and the output is the written transcript. The transcript is a rich output because it contains a sequence of words, not just one label or one number.

Key points:

  • Input: audio recording.
  • Output: written transcript.
  • The output is more detailed than a single label or score.

Rubric: Award points for identifying audio as the input, transcript as the output, and explaining that the output is richer than a single number or class label.

0

1

Updated 2026-08-12

Contributors are:

Who are from:

Tags

Machine Learning

Deep Learning

Supervised Learning

Dive into Deep Learning @ D2L

Data Science

Machine Learning Strategy

Machine Learning Yearning @ DeepLearning.AI