You are tasked with preparing a dataset for a human feedback-based model tuning process. The initial dataset consists only of user prompts. Arrange the following actions into the correct chronological sequence to create the initial set of data for human evaluation.
0
1
Tags
Ch.4 Alignment - Foundations of Large Language Models
Foundations of Large Language Models
Foundations of Large Language Models Course
Computing Sciences
Comprehension in Revised Bloom's Taxonomy
Cognitive Psychology
Psychology
Social Science
Empirical Science
Science
Related
Comparison of Annotation Methods for Human Feedback in RLHF
A development team is refining a large language model to be more helpful and safe using feedback from human evaluators. For the prompt, 'Explain the water cycle for a 10-year-old,' the model generates four different responses:
- 'Rain falls, flows to the sea, evaporates into clouds, and rains again.'
- 'Imagine water goes on a big trip! It falls from clouds as rain, runs into rivers, then the sun warms it up until it floats back into the sky to make new clouds.'
- 'The water cycle describes t
Evaluating Output Sets for Human Feedback
Formulating the Loss Function for Policy Learning in RLHF
You are tasked with preparing a dataset for a human feedback-based model tuning process. The initial dataset consists only of user prompts. Arrange the following actions into the correct chronological sequence to create the initial set of data for human evaluation.