Multilingual Benchmark Translation for LLM Evaluation
To evaluate language model capabilities across diverse linguistic contexts where native test sets are lacking, comprehensive English-language benchmarks—such as the 57-subject multiple-choice MMLU suite—are translated into target languages using automated machine translation systems like Azure Translate. This approach facilitates standardized cross-lingual evaluation across a wide spectrum of languages, spanning high-resource languages with hundreds of millions of speakers down to low-resource languages.
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.2 Model Scaling and Capability Evaluation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Multilingual Language Understanding on MMLU - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
Multilingual Benchmark Translation for LLM Evaluation
GPT-4 Multilingual Performance on MMLU
GPT-4
MMLU Benchmark
Challenges of Multilingual LLMs for Low-Resource Languages
Example of an MMLU Question (Abstract Algebra)
Match each MMLU benchmark component to its role or definition described in the course content.
Explain why this researcher's proposed protocol conflicts with the format mandated by the MMLU benchmark, and identify what the model is required to produce instead.
According to the course content, which broad category of tasks does the MMLU benchmark organize into a question-answering format?
What full name does the acronym "MMLU" represent?
Multilingual Benchmark Translation for LLM Evaluation
Learn After
Which automated machine translation system is used to translate comprehensive English-language benchmarks into target languages for cross-lingual evaluation?
Translating comprehensive benchmarks enables standardized cross-lingual evaluation across both high-resource and low-resource languages.
Under what circumstance are comprehensive English-language benchmarks translated into target languages for language model evaluation?
GPT-4 Multilingual Performance on MMLU