Multiple Choice

A team is refining a text-generation model. The final probability of a generated text sequence is proportional to the product of its probability from an initial base model and an exponentiated reward score. The reward's influence is controlled by a scaling parameter, β, in the exponent, where a smaller β gives the reward more weight. The team observes that when they significantly decrease the value of β, the model's outputs become more repetitive and sometimes nonsensical, even though they achie

0

1

Updated 2025-09-26

Contributors are:

Who are from:

Tags

Ch.4 Alignment - Foundations of Large Language Models

Foundations of Large Language Models

Foundations of Large Language Models Course

Computing Sciences

Analysis in Bloom's Taxonomy

Cognitive Psychology

Psychology

Social Science

Empirical Science

Science