OpenAI GPT-5.6 System Card - Challenging Prompts - Non-Violent Illicit Behavior: leaderboard

Metric: not_unsafe rate (%). Source: deploymentsafety.openai.com. 7 models tracked.

Top models

#ModelScore
1GPT-5.4 (Thinking)100
2GPT-5.6 Luna99.3
3GPT-5.2 (Thinking)99.3
4GPT-5.6 Terra99
5GPT-5.1 (Thinking)99
6GPT-5.598.7
7GPT-5.6 Sol98.7

Interactive version: theaggregate.ai/benchmark?slug=openai-gpt-5-6-system-card-challenging-prompts-non-violent-illicit-behavior · How It Works · Data refreshed daily, snapshot 2026-09-19.