Multiple Choice

A student model is trained to mimic a teacher model by minimizing the following loss function, which measures the dissimilarity between their output probability distributions for a given input:

Loss=yPrt(y)logPrθs(y)\text{Loss} = -\sum_{\mathbf{y}} \text{Pr}^t(\mathbf{y}) \log \text{Pr}_{\theta}^s(\mathbf{y})

In this formula, Prt(y)\text{Pr}^t(\mathbf{y}) is the teacher's probability for an output sequence y\mathbf{y}, Prθs(y)\text{Pr}_{\theta}^s(\mathbf{y}) is the student's probability, and the summation is over all

0

1

Updated 2025-09-26

Contributors are:

Who are from:

Tags

Ch.3 Prompting - Foundations of Large Language Models

Foundations of Large Language Models

Foundations of Large Language Models Course

Computing Sciences

Analysis in Bloom's Taxonomy

Cognitive Psychology

Psychology

Social Science

Empirical Science

Science