Contamination as a Non-Substantive Confounder in GPT-4 Exam Evaluation
Across comprehensive evaluations on standardized academic and professional exams, pre-training data contamination did not act as a substantive confounding factor for GPT-4's performance. The calculated degradation between non-contaminated and contaminated questions was generally minor and occurred approximately as often with a positive sign as with a negative sign, demonstrating that prior exposure to exam questions in the training corpus did not artificially drive the model's high test scores.
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.2 Model Scaling and Capability Evaluation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Test Set Contamination Analysis - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
Measuring Test Set Contamination via Substring Matching
Performance Degradation Metric for Contamination Evaluation
Contamination as a Non-Substantive Confounder in GPT-4 Exam Evaluation
GPT-4 Performance on Academic and Professional Exams
How is performance degradation calculated when evaluating test set contamination in language models?
In test set contamination analysis, non-contaminated questions represent leaked items, while contaminated questions represent unseen items.
In the context of test set contamination analysis, what does the performance degradation metric quantify?
A language model achieves a score of 62% on the non-contaminated questions of an evaluation benchmark and 88% on the contaminated questions of the same benchmark.
- Calculate the performance degradation metric for this evaluation.
- Interpret what this specific result indicates regarding the model's reported capability metrics.
Contamination as a Non-Substantive Confounder in GPT-4 Exam Evaluation
Learn After
What did comprehensive evaluations of GPT-4 on standardized academic and professional exams reveal regarding the effect of pre-training data contamination?
In the evaluation of GPT-4's exam results, how did the calculated degradation between non-contaminated and contaminated questions behave in terms of both its magnitude and directional sign?