How is performance degradation calculated when evaluating test set contamination in language models?
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.2 Model Scaling and Capability Evaluation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Test Set Contamination Analysis - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
How is performance degradation calculated when evaluating test set contamination in language models?
In test set contamination analysis, non-contaminated questions represent leaked items, while contaminated questions represent unseen items.
In the context of test set contamination analysis, what does the performance degradation metric quantify?
A language model achieves a score of 62% on the non-contaminated questions of an evaluation benchmark and 88% on the contaminated questions of the same benchmark.
- Calculate the performance degradation metric for this evaluation.
- Interpret what this specific result indicates regarding the model's reported capability metrics.
Contamination as a Non-Substantive Confounder in GPT-4 Exam Evaluation