Multiple Choice

In a reinforcement learning process, a new policy defined by parameters θ is evaluated using an objective function that relies on data from a reference policy with parameters θ_ref. The objective function is:

J(θ) = E_{τ ~ π_{θ_ref}} [ (Pr_θ(τ) / Pr_{θ_ref}(τ)) * R(τ) ]

Where τ is a trajectory, Pr(τ) is the probability of that trajectory, R(τ) is its total reward, and E_{τ ~ π_{θ_ref}} denotes the expected value over trajectories from the reference policy.

What does this objective functi

0

1

Updated 2025-09-28

Contributors are:

Who are from:

Tags

Ch.4 Alignment - Foundations of Large Language Models

Foundations of Large Language Models

Foundations of Large Language Models Course

Computing Sciences

Application in Bloom's Taxonomy

Cognitive Psychology

Psychology

Social Science

Empirical Science

Science