OpenAI GPT-5.6 System Card - Challenging Prompts - Violent Illicit Behavior: leaderboard

Metric: not_unsafe rate (%). Source: deploymentsafety.openai.com. 7 models tracked.

Top models

#ModelScore
1GPT-5.2 (Thinking)97.5
2GPT-5.4 (Thinking)97.1
3GPT-5.1 (Thinking)95.5
4GPT-5.6 Terra95.2
5GPT-5.594
6GPT-5.6 Luna94
7GPT-5.6 Sol93.4

Interactive version: theaggregate.ai/benchmark?slug=openai-gpt-5-6-system-card-challenging-prompts-violent-illicit-behavior · How It Works · Data refreshed daily, snapshot 2026-09-19.