Learn Before
Reasoning Tasks as Question Answering
Multiple-Choice Question Answering
Multilingual Language Understanding on MMLU - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Model Calibration and Confidence Degradation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
MMLU Benchmark
The MMLU (Massive Multitask Language Understanding) benchmark is a prominent example of how complex reasoning tasks are structured in a question-answering format. Each problem within this benchmark is presented as a multiple-choice question, requiring a Large Language Model to choose the single correct answer from a list of options.
0
1
Tags
Ch.3 Prompting - Foundations of Large Language Models
Foundations of Large Language Models
Computing Sciences
Foundations of Large Language Models Course
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.2 Model Scaling and Capability Evaluation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Multilingual Language Understanding on MMLU - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.3 Model Alignment and Safety - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Model Calibration and Confidence Degradation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
MMLU Benchmark
A product development team is using a large language model to check if a new product concept aligns with their company's core principles. Their initial prompt, "Analyze if our new 'Smart-Mug' concept is consistent with our principles of 'sustainability,' 'simplicity,' and 'affordability'," yields vague and unhelpful responses. How could this reasoning task be most effectively restructured into a question-answering format to guide the model toward a more structured and deductive output?
Improving a Data Analysis Prompt
Reframing a Research Query
MMLU Benchmark
A team of engineers is evaluating a new language model's reasoning capabilities. They use an assessment method where the model must choose the single correct answer from a set of provided options for each question. Which of the following represents a primary limitation of this evaluation method for gauging the model's genuine comprehension?
AI Tutor Design Strategy
Designing a Challenging Multiple-Choice Question for a Language Model
Example of a Sentence-First Prompt for Grammaticality Judgment with Answer Options
Multilingual Benchmark Translation for LLM Evaluation
GPT-4 Multilingual Performance on MMLU
GPT-4
MMLU Benchmark
Challenges of Multilingual LLMs for Low-Resource Languages
Degradation of Model Calibration from Post-Training Alignment
MMLU Benchmark
Learn After
Example of an MMLU Question (Abstract Algebra)
Match each MMLU benchmark component to its role or definition described in the course content.
Explain why this researcher's proposed protocol conflicts with the format mandated by the MMLU benchmark, and identify what the model is required to produce instead.
According to the course content, which broad category of tasks does the MMLU benchmark organize into a question-answering format?
What full name does the acronym "MMLU" represent?
Multilingual Benchmark Translation for LLM Evaluation