Concept icon
Concept

GPT-4 Performance on Academic and Professional Exams

GPT-4 was evaluated across a wide range of academic and professional exams designed for humans, achieving human-level performance on the majority of them. On a simulated Uniform Bar Examination, GPT-4 scores around the 90th percentile (top 10% of test takers), contrasting sharply with GPT-3.5, which scored in the bottom 10%. Across diverse standardized tests—including the LSAT, SAT, GRE, and various Advanced Placement (AP) exams—GPT-4 significantly outperforms prior models such as GPT-3.5.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Ch.2 Model Scaling and Capability Evaluation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Academic and Professional Exam Benchmarks - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Test Set Contamination Analysis - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor