True/False

Consider the following equation that defines a target policy πθ\pi_{\theta} based on a reference policy πθref\pi_{\theta_{\text{ref}}}, a reward function r(x,y)r(\mathbf{x}, \mathbf{y}), a positive scaling parameter β\beta, and a normalization term Z(x)Z(\mathbf{x}): πθ(yx)=πθref(yx)exp(1βr(x,y))Z(x)\pi_{\theta}(\mathbf{y}|\mathbf{x}) = \frac{\pi_{\theta_{\text{ref}}}(\mathbf{y}|\mathbf{x}) \exp(\frac{1}{\beta}r(\mathbf{x}, \mathbf{y}))}{Z(\mathbf{x})} True or False: If the reward function r(x,y)r(\mathbf{x}, \mathbf{y}) is equal to

0

1

Updated 2025-10-04

Contributors are:

Who are from:

Tags

Ch.4 Alignment - Foundations of Large Language Models

Foundations of Large Language Models

Foundations of Large Language Models Course

Computing Sciences

Analysis in Bloom's Taxonomy

Cognitive Psychology

Psychology

Social Science

Empirical Science

Science