Learn Before
Model Calibration and Confidence Degradation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Model Calibration and Post-Training Effects - Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor
RLHF Effects on Capability and Calibration - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Degradation of Model Calibration from Post-Training Alignment
Post-training alignment (such as PPO reinforcement learning) can significantly degrade the calibration of large language models. In pre-trained models such as GPT-4, predicted probabilities (logprobs) across multiple-choice options track actual task accuracy closely, yielding near-perfect calibration with an Expected Calibration Error (ECE) of $0.007on benchmarks like MMLU. However, subsequent post-training alignment impairs this correspondence, substantially inflating calibration error to an ECE of $0.074 and causing the model's confidence to diverge from its true probability of being correct.
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.3 Model Alignment and Safety - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Model Calibration and Confidence Degradation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor
Ch.2 Post-Training Analysis and Safety - Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor
Model Calibration and Post-Training Effects - Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor
Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Ch.2 Post-Training Alignment and Calibration - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
RLHF Effects on Capability and Calibration - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Related
Degradation of Model Calibration from Post-Training Alignment
MMLU Benchmark
Degradation of Model Calibration from Post-Training Alignment
Degradation of Model Calibration from Post-Training Alignment
Origin of GPT-4 Exam Capabilities in Pre-training
Evaluation Asymmetry in Base and RLHF Free-Response Comparison
Learn After
Match each stage, metric value, or procedure to its role in large language model calibration on the MMLU benchmark.
Order the progression of model calibration and confidence behavior from the pre-trained state through post-training alignment.
Explain how the post-training alignment impacted the model's calibration and what the increase in ECE signifies regarding the model's confidence versus its accuracy.
In pre-trained GPT-4, what relationship is observed between the model's predicted probabilities (logprobs) across multiple-choice options and its actual task accuracy on benchmarks like MMLU?
Based on benchmark evaluations on MMLU, how does post-training alignment quantitatively alter GPT-4's Expected Calibration Error (ECE)?