Evaluation Asymmetry in Base and RLHF Free-Response Comparison
Evaluating pre-trained base models against post-RLHF models on an equal footing is difficult for free-response tasks compared to multiple-choice benchmarks. Free-response answer generation and sampling methodologies rely heavily on the model's capacity to follow open-ended instructions accurately, an attribute developed during instruction-tuning and RLHF post-training that raw pre-trained base models inherently lack.
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.3 Model Alignment and Safety - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Impact of RLHF on Model Capability - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
Origin of GPT-4 Exam Capabilities in Pre-training
Evaluation Asymmetry in Base and RLHF Free-Response Comparison
When evaluated on multiple-choice standardized exam sections, what average score is achieved by the base pre-trained GPT-4 model?
Empirical evaluation shows that GPT-4's standardized exam benchmark capabilities originate primarily from post-training alignment rather than pre-training.
Based on empirical evaluations on standardized exam benchmarks, what effect does reinforcement learning from human feedback (RLHF) have on the fundamental capabilities acquired by GPT-4 during pre-training?
In empirical evaluations of GPT-4, which finding provides evidence that standardized exam performance originates primarily in pre-training rather than post-training alignment?
On what specific format of standardized exam sections were the base pre-trained and post-RLHF GPT-4 models evaluated to compare their benchmark capabilities?
Evaluation Asymmetry in Base and RLHF Free-Response Comparison
Learn After
When attempting to evaluate pre-trained base models against post-RLHF models on an equal footing, which format presents fewer evaluation difficulties than free-response tasks?
Raw pre-trained base models inherently possess the capacity to accurately follow open-ended instructions.
Which post-training processes develop a model's capacity to follow open-ended instructions accurately?
Explain why answer generation and sampling methodologies create an evaluation asymmetry when comparing pre-trained base models to post-RLHF models on free-response tasks.