Concept icon
Concept

Safety Metric Improvements in GPT-4

Safety mitigations and post-training interventions substantially improve GPT-4's safety performance relative to GPT-3.5. Specifically, GPT-4 exhibits an $82%decrease in the tendency to respond to requests for disallowed content. On sensitive requests, such as those regarding medical advice or self-harm, GPT-4 complies with policy guidelines $29% more often. Additionally, on the RealToxicityPrompts benchmark, GPT-4 produces toxic generations only $0.73%of the time, compared to $6.48% for GPT-3.5.

0

1

Concept icon
Updated 2026-09-11

Tags

Prep Sessions

Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor

Ch.2 Post-Training Analysis and Safety - Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor

Safety Metrics and Refusal Behavior - Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor