Learn Before
Multilingual Language Understanding on MMLU - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Multilingual Benchmark Translation for LLM Evaluation
Multilingual Generalization on MMLU - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
GPT-4 Multilingual Performance on MMLU
On translated versions of the MMLU benchmark, GPT-4 demonstrates robust cross-lingual understanding by outperforming the English-language benchmark scores of preceding state-of-the-art models (such as Chinchilla and PaLM) across the vast majority of evaluated languages. Notably, GPT-4 surpasses prior English-language models even in low-resource languages with limited pre-training data availability, such as Latvian, Welsh, and Swahili.
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.2 Model Scaling and Capability Evaluation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Multilingual Language Understanding on MMLU - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Ch.1 Foundation Model Capabilities and Benchmarking - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Multilingual Generalization on MMLU - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Related
Multilingual Benchmark Translation for LLM Evaluation
GPT-4 Multilingual Performance on MMLU
GPT-4
MMLU Benchmark
Challenges of Multilingual LLMs for Low-Resource Languages
Which automated machine translation system is used to translate comprehensive English-language benchmarks into target languages for cross-lingual evaluation?
Translating comprehensive benchmarks enables standardized cross-lingual evaluation across both high-resource and low-resource languages.
Under what circumstance are comprehensive English-language benchmarks translated into target languages for language model evaluation?
GPT-4 Multilingual Performance on MMLU
GPT-4 Multilingual Performance on MMLU
Learn After
Match each component of the MMLU evaluation to its role or description in assessing GPT-4:
Assess whether the reviewer's hypothesis is supported by the evaluation results and explain how GPT-4's performance in these specific languages compares to the English scores of prior models.
On translated versions of the MMLU benchmark, which models' English-language benchmark scores does GPT-4 outperform across the vast majority of evaluated languages?
Why is GPT-4's performance on languages such as Latvian, Welsh, and Swahili considered particularly notable on translated MMLU evaluations?