Multiple Choice

A development team is training a large language model to be a helpful assistant. Their process involves two stages:

  1. They train a 'scoring model' on a dataset of human-ranked conversations. The goal of this scoring model is to predict which of two responses a human would prefer, assigning a numerical score.
  2. They then use this scoring model to automatically provide feedback to the main language model, rewarding it for generating responses that receive a high score.

After extensive training

0

1

Updated 2025-09-26

Contributors are:

Who are from:

Tags

Ch.2 Generative Models - Foundations of Large Language Models

Foundations of Large Language Models

Foundations of Large Language Models Course

Computing Sciences

Ch.4 Alignment - Foundations of Large Language Models

Evaluation in Bloom's Taxonomy

Cognitive Psychology

Psychology

Social Science

Empirical Science

Science