Short Answer

A reward model in an RLHF framework processes an interaction where a user submits the prompt "Summarize the plot of Hamlet" and receives the completion "Hamlet is a tragedy by William Shakespeare about Prince Hamlet seeking revenge against his uncle Claudius," yielding a result of 2.5. Based on the structure of an RLHF reward model, identify the two token sequence inputs (x and y) and specify what the resulting output value of 2.5 represents.

0

1

Updated 2026-09-07

Contributors are:

Who are from:

Tags

Ch.4 Alignment - Foundations of Large Language Models

Foundations of Large Language Models

Foundations of Large Language Models Course

Computing Sciences

Application in Bloom's Taxonomy

Cognitive Psychology

Psychology

Social Science

Empirical Science

Science

Prep Sessions

Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Ch.3 Model Alignment and Safety - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Model-Assisted Safety and Rule-Based Reward Models - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

OpenStax Psychology (2nd ed.) Textbook