Concept icon
Concept

Evaluation Asymmetry in Base and RLHF Free-Response Comparison

Evaluating pre-trained base models against post-RLHF models on an equal footing is difficult for free-response tasks compared to multiple-choice benchmarks. Free-response answer generation and sampling methodologies rely heavily on the model's capacity to follow open-ended instructions accurately, an attribute developed during instruction-tuning and RLHF post-training that raw pre-trained base models inherently lack.

0

1

Concept icon
Updated 2026-09-11

Tags

Prep Sessions

Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Ch.3 Model Alignment and Safety - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Impact of RLHF on Model Capability - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor

Ch.2 Post-Training Alignment and Calibration - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor

RLHF Effects on Capability and Calibration - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor