1Cademy - A research team is refining a language models ability to be helpful and harmless. They use two distinct datasets for this process. Dataset 1 contains prompts, each paired with a single, meticulously crafted, ideal response. Dataset 2 contains prompts, each paired with two different model-generated responses, along with a label indicating which of the two responses a human preferred. Which statement best distinguishes the fundamental optimization objective when training on Dataset 1 versus Dataset 2?

Learn Before

Comparison of Objectives: Supervised Fine-Tuning vs. RLHF

Multiple Choice

A research team is refining a language model's ability to be helpful and harmless. They use two distinct datasets for this process. Dataset 1 contains prompts, each paired with a single, meticulously crafted, ideal response. Dataset 2 contains prompts, each paired with two different model-generated responses, along with a label indicating which of the two responses a human preferred. Which statement best distinguishes the fundamental optimization objective when training on Dataset 1 versus Dataset 2?

Updated 2025-10-06

Contributors are:

Who are from:

Learn Before

Related