Match each standardized exam benchmark metric with its corresponding empirical result for GPT-4.
0
1
Tags
Prep Sessions
Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Ch.2 Post-Training Alignment and Calibration - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
RLHF Effects on Capability and Calibration - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Related
Empirical evaluation shows that GPT-4's standardized exam benchmark capabilities originate primarily from post-training alignment rather than pre-training.
In empirical evaluations of GPT-4, which finding provides evidence that standardized exam performance originates primarily in pre-training rather than post-training alignment?
On what specific format of standardized exam sections were the base pre-trained and post-RLHF GPT-4 models evaluated to compare their benchmark capabilities?
Evaluation Asymmetry in Base and RLHF Free-Response Comparison
Match each standardized exam benchmark metric with its corresponding empirical result for GPT-4.
Which section format was used in the empirical evaluations comparing the standardized benchmark performance of base and post-RLHF GPT-4?
In which stage of model development does GPT-4's performance across standardized exam benchmarks primarily originate?