In the evaluation of GPT-4's exam results, how did the calculated degradation between non-contaminated and contaminated questions behave in terms of both its magnitude and directional sign?
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.2 Model Scaling and Capability Evaluation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Test Set Contamination Analysis - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
What did comprehensive evaluations of GPT-4 on standardized academic and professional exams reveal regarding the effect of pre-training data contamination?
In the evaluation of GPT-4's exam results, how did the calculated degradation between non-contaminated and contaminated questions behave in terms of both its magnitude and directional sign?