logo
How it worksCoursesResearch CommunitiesBenefitsAbout Us
Schedule Demo
Learn Before
  • Examples of Constant Segment-Based Reward Functions

True/False

A reward function that assigns a constant positive value (e.g., +1) to every segment of a generated text is an effective method for training a model to differentiate between well-written and poorly-written segments.

0

1

Updated 2025-10-10

Contributors are:

G
Gemini AI
🏆 2

Who are from:

G
Google
🏆 2

Tags

Ch.4 Alignment - Foundations of Large Language Models

Foundations of Large Language Models

Computing Sciences

Foundations of Large Language Models Course

Analysis in Bloom's Taxonomy

Cognitive Psychology

Psychology

Social Science

Empirical Science

Science

Related
  • A machine learning team is training a model to generate creative stories. They implement a reward mechanism where every segment of a generated story is assigned a score of exactly +1, irrespective of the segment's content, the initial prompt, or the rest of the story. Which of the following outcomes is the most likely consequence of this reward strategy?

  • Evaluating a Chatbot's Reward Function

  • A reward function that assigns a constant positive value (e.g., +1) to every segment of a generated text is an effective method for training a model to differentiate between well-written and poorly-written segments.

  • Formula for a Negative Reward Function r(x,y,yˉ)r(\mathbf{x}, \mathbf{y}, \bar{\mathbf{y}})r(x,y,yˉ​)

  • Formula for a Positive Reward Function r(x,y,yˉ)r(\mathbf{x}, \mathbf{y}, \bar{\mathbf{y}})r(x,y,yˉ​)

logo 1cademy1Cademy

Optimize Scalable Learning and Teaching

How it worksCoursesResearch CommunitiesBenefitsAbout UsAll Courses
TermsPrivacyCookieGDPR

Contact Us

iman@honor.education

Follow Us




© 1Cademy 2026

We're committed to OpenSource on

Github