Learn Before
Academic and Professional Exam Benchmarks - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Impact of RLHF on Model Capability - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
GPT-4 Performance on Academic and Professional Exams
Origin of GPT-4 Exam Capabilities in Pre-training
Empirical evaluation shows that GPT-4's performance across standardized exam benchmarks originates primarily in the pre-training stage rather than post-training alignment. When tested on multiple-choice exam sections, the base pre-trained GPT-4 model attains an average score of 73.7% while the post-RLHF model attains 74.0%, indicating that reinforcement learning from human feedback does not substantially alter the fundamental capabilities acquired during pre-training.
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.2 Model Scaling and Capability Evaluation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Academic and Professional Exam Benchmarks - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.3 Model Alignment and Safety - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Impact of RLHF on Model Capability - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
GPT-4 Performance on Academic and Professional Exams
Origin of GPT-4 Exam Capabilities in Pre-training
GPT-4
Origin of GPT-4 Exam Capabilities in Pre-training
Evaluation Asymmetry in Base and RLHF Free-Response Comparison
How did GPT-4's performance compare to GPT-3.5 on the simulated Uniform Bar Examination?
Across the majority of academic and professional exams designed for humans, what benchmark level of performance did GPT-4 achieve?
The academic and professional examinations used to evaluate GPT-4 share which design characteristic?
Aside from the Uniform Bar Examination, which four standardized examination programs are explicitly cited as part of GPT-4's evaluation?
Origin of GPT-4 Exam Capabilities in Pre-training
Learn After
When evaluated on multiple-choice standardized exam sections, what average score is achieved by the base pre-trained GPT-4 model?
Empirical evaluation shows that GPT-4's standardized exam benchmark capabilities originate primarily from post-training alignment rather than pre-training.
Based on empirical evaluations on standardized exam benchmarks, what effect does reinforcement learning from human feedback (RLHF) have on the fundamental capabilities acquired by GPT-4 during pre-training?
In empirical evaluations of GPT-4, which finding provides evidence that standardized exam performance originates primarily in pre-training rather than post-training alignment?
On what specific format of standardized exam sections were the base pre-trained and post-RLHF GPT-4 models evaluated to compare their benchmark capabilities?
Evaluation Asymmetry in Base and RLHF Free-Response Comparison