GPT-4 Multilingual Performance on MMLU
On translated versions of the MMLU benchmark, GPT-4 demonstrates robust cross-lingual understanding by outperforming the English-language benchmark scores of preceding state-of-the-art models (such as Chinchilla and PaLM) across the vast majority of evaluated languages. Notably, GPT-4 surpasses prior English-language models even in low-resource languages with limited pre-training data availability, such as Latvian, Welsh, and Swahili.
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.2 Model Scaling and Capability Evaluation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Multilingual Language Understanding on MMLU - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
Multilingual Benchmark Translation for LLM Evaluation
GPT-4 Multilingual Performance on MMLU
GPT-4
MMLU Benchmark
Challenges of Multilingual LLMs for Low-Resource Languages
Which automated machine translation system is used to translate comprehensive English-language benchmarks into target languages for cross-lingual evaluation?
Translating comprehensive benchmarks enables standardized cross-lingual evaluation across both high-resource and low-resource languages.
Under what circumstance are comprehensive English-language benchmarks translated into target languages for language model evaluation?
GPT-4 Multilingual Performance on MMLU