Multiple Choice

A reward model is being trained with a loss function that includes a regularization component. This component adds a penalty proportional to (r(x,ya)+r(x,yb))2(r(\mathbf{x}, \mathbf{y}_a) + r(\mathbf{x}, \mathbf{y}_b))^2 for a given input x\mathbf{x} and a pair of responses (ya,yb)(\mathbf{y}_a, \mathbf{y}_b). The goal of this penalty is to prevent reward scores from becoming excessively large. Consider two scenarios for the reward scores assigned to a pair of responses:

  • Scenario 1: $r(\mathbf{x}, \mathbf{y}_a)

0

1

Updated 2025-10-04

Contributors are:

Who are from:

Tags

Ch.4 Alignment - Foundations of Large Language Models

Foundations of Large Language Models

Computing Sciences

Foundations of Large Language Models Course

Analysis in Bloom's Taxonomy

Cognitive Psychology

Psychology

Social Science

Empirical Science

Science