Comparison

Safety Metric Improvements in GPT-4

Safety mitigations and post-training interventions substantially improve GPT-4's safety performance relative to GPT-3.5. GPT-4 shows an 82%82\% decrease in its tendency to respond to requests for disallowed content and follows policy guidelines on sensitive requests, such as medical advice and self-harm, 29%29\% more often. On the RealToxicityPrompts benchmark, GPT-4 produces toxic generations 0.73% of the time, compared with 6.48% for GPT-3.5.

0

1

Updated 2026-09-12

Tags

Prep Sessions

Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor

Ch.2 Post-Training Analysis and Safety - Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor

Safety Metrics and Refusal Behavior - Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor