GPT-4 Performance on Academic and Professional Exams
GPT-4 was evaluated across a wide range of academic and professional exams designed for humans, achieving human-level performance on the majority of them. On a simulated Uniform Bar Examination, GPT-4 scores around the 90th percentile (top 10% of test takers), contrasting sharply with GPT-3.5, which scored in the bottom 10%. Across diverse standardized tests—including the LSAT, SAT, GRE, and various Advanced Placement (AP) exams—GPT-4 significantly outperforms prior models such as GPT-3.5.
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.2 Model Scaling and Capability Evaluation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Academic and Professional Exam Benchmarks - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Test Set Contamination Analysis - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
GPT-4 Performance on Academic and Professional Exams
Origin of GPT-4 Exam Capabilities in Pre-training
GPT-4
Measuring Test Set Contamination via Substring Matching
Performance Degradation Metric for Contamination Evaluation
Contamination as a Non-Substantive Confounder in GPT-4 Exam Evaluation
GPT-4 Performance on Academic and Professional Exams
Which statement accurately describes the disclosure of GPT-4's technical specifications?
Discuss the operational implications of GPT-4's multimodal architecture compared to its text-only predecessors in the GPT series. In your response, identify the specific input and output modalities that define GPT-4, and explain how the ability to process multiple data types expands its functional capabilities beyond earlier text-only models.
What type of output does GPT-4 generate when processing inputs?
According to the text, which model is GPT-4 the direct successor to?
Previous architectures in the GPT series prior to GPT-4 were text-only models.
According to the text, what specific scale descriptor is applied to the GPT-4 model?
GPT-4 Performance on Academic and Professional Exams
Multimodal Input Processing in GPT-4
Learn After
How did GPT-4's performance compare to GPT-3.5 on the simulated Uniform Bar Examination?
Across the majority of academic and professional exams designed for humans, what benchmark level of performance did GPT-4 achieve?
The academic and professional examinations used to evaluate GPT-4 share which design characteristic?
Aside from the Uniform Bar Examination, which four standardized examination programs are explicitly cited as part of GPT-4's evaluation?
Origin of GPT-4 Exam Capabilities in Pre-training