Short Answer

In the evaluation of GPT-4's exam results, how did the calculated degradation between non-contaminated and contaminated questions behave in terms of both its magnitude and directional sign?

0

1

Updated 2026-09-07

Tags

Prep Sessions

Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Ch.2 Model Scaling and Capability Evaluation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Test Set Contamination Analysis - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor