Safety Metric Improvements in GPT-4
Safety mitigations and post-training interventions substantially improve GPT-4's safety performance relative to GPT-3.5. GPT-4 shows an decrease in its tendency to respond to requests for disallowed content and follows policy guidelines on sensitive requests, such as medical advice and self-harm, more often. On the RealToxicityPrompts benchmark, GPT-4 produces toxic generations 0.73% of the time, compared with 6.48% for GPT-3.5.
0
1
Contributors are:
Who are from:
Tags
Prep Sessions
Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor
Ch.2 Post-Training Analysis and Safety - Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor
Safety Metrics and Refusal Behavior - Frontier Model Dynamics: Scaling Laws, Calibration, and Post-Training Alignment @ University of Michigan - Ann Arbor
Learn After
According to the course content, which interventions substantially improved GPT-4's safety performance compared to GPT-3.5?
True or False: On the RealToxicityPrompts benchmark, GPT-4 produces toxic generations at a higher rate than GPT-3.5.
By what percentage does GPT-4 decrease its tendency to respond to requests for disallowed content compared to GPT-3.5?
Synthesize the safety improvements of GPT-4 compared to GPT-3.5 across all three evaluated safety dimensions: disallowed content, sensitive requests, and toxic generation benchmarks. Detail the specific performance metrics achieved in each area.