Essay

A language model achieves a score of 62% on the non-contaminated questions of an evaluation benchmark and 88% on the contaminated questions of the same benchmark.

  1. Calculate the performance degradation metric for this evaluation.
  2. Interpret what this specific result indicates regarding the model's reported capability metrics.

0

1

Updated 2026-09-11

Tags

Prep Sessions

Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Ch.2 Model Scaling and Capability Evaluation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Test Set Contamination Analysis - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor

Ch.1 Foundation Model Capabilities and Benchmarking - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor

Benchmark Contamination Analysis - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor