Safety Metric Improvements in GPT-4
Safety mitigations and post-training interventions substantially improve GPT-4's safety performance relative to GPT-3.5. Specifically, GPT-4 exhibits an $82%decrease in the tendency to respond to requests for disallowed content. On sensitive requests, such as those regarding medical advice or self-harm, GPT-4 complies with policy guidelines $29% more often. Additionally, on the RealToxicityPrompts benchmark, GPT-4 produces toxic generations only $0.73%of the time, compared to $6.48% for GPT-3.5.
0
1
Tags
Prep Sessions
Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor
Ch.2 Post-Training Analysis and Safety - Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor
Safety Metrics and Refusal Behavior - Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor
Learn After
According to the course content, which interventions substantially improved GPT-4's safety performance compared to GPT-3.5?
True or False: On the RealToxicityPrompts benchmark, GPT-4 produces toxic generations at a higher rate than GPT-3.5.
By what percentage does GPT-4 decrease its tendency to respond to requests for disallowed content compared to GPT-3.5?
Synthesize the safety improvements of GPT-4 compared to GPT-3.5 across all three evaluated safety dimensions: disallowed content, sensitive requests, and toxic generation benchmarks. Detail the specific performance metrics achieved in each area.