Learn Before
On translated versions of the MMLU benchmark, which models' English-language benchmark scores does GPT-4 outperform across the vast majority of evaluated languages?
0
1
Tags
Prep Sessions
Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Ch.1 Foundation Model Capabilities and Benchmarking - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Multilingual Generalization on MMLU - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Related
Match each component of the MMLU evaluation to its role or description in assessing GPT-4:
Assess whether the reviewer's hypothesis is supported by the evaluation results and explain how GPT-4's performance in these specific languages compares to the English scores of prior models.
On translated versions of the MMLU benchmark, which models' English-language benchmark scores does GPT-4 outperform across the vast majority of evaluated languages?
Why is GPT-4's performance on languages such as Latvian, Welsh, and Swahili considered particularly notable on translated MMLU evaluations?