Multiple Choice

In a particular policy optimization framework, the target policy, denoted as πθ(yx)\pi_{\theta}(\mathbf{y}|\mathbf{x}), is determined by the following relationship involving a reference policy πθref\pi_{\theta_{\text{ref}}}, a reward function r(x,y)r(\mathbf{x}, \mathbf{y}), a positive temperature parameter β\beta, and a normalization term Z(x)Z(\mathbf{x}): $$ \pi_{\theta}(\mathbf{y}|\mathbf{x}) = \frac{\pi_{\theta_{\text{ref}}}(\mathbf{y}|\mathbf{x}) \exp(\frac{1}{\beta}r(\mathbf{x}, \mathbf{y}))}{Z(\math

0

1

Updated 2025-09-28

Contributors are:

Who are from:

Tags

Ch.4 Alignment - Foundations of Large Language Models

Foundations of Large Language Models

Foundations of Large Language Models Course

Computing Sciences

Analysis in Bloom's Taxonomy

Cognitive Psychology

Psychology

Social Science

Empirical Science

Science