GPT-5.6 Luna (August 2026): benchmark results

Provider: OpenAI. Released 2026-07-09. Access: API.

Unified ELO 1875 ± 41, rank #54 of 1516 rated models, from 44 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
OpenAI GPT-5.6 August Update System Card - Adversarial User Simulations - Self-Harm91.1not_unsafe message rate (%)100
OpenAI GPT-5.6 August Update System Card - Challenging Prompts - Gore86.5not_unsafe rate (%)100
OpenAI GPT-5.6 August Update System Card - Challenging Prompts - Sexual97.4not_unsafe rate (%)100
OpenAI GPT-5.6 August Update System Card - Challenging Prompts - Sexual/Minors96.6not_unsafe rate (%)100
OpenAI GPT-5.6 August Update System Card - U18 - Eating Disorders81Score (%)100
OpenAI GPT-5.6 August Update System Card - U18 - Sexual Content98.4Score (%)100
OpenAI MentalHealthBench - Behavior - Interpretation69.91Axis score, normalized to its possible range (%)78.6
OpenAI GPT-5.6 August Update System Card - U18 - Emotional Reliance92.7Score (%)75
OpenAI MentalHealthBench - Behavior - Communication67.06Axis score, normalized to its possible range (%)71.4
OpenAI MentalHealthBench - Behavior - Reality Testing76.85Axis score, normalized to its possible range (%)71.4
OpenAI MentalHealthBench - Behavior - Empathy and Support82.21Axis score, normalized to its possible range (%)67.9
OpenAI GPT-5.6 August Update System Card - Challenging Prompts - Non-Violent Illicit Behavior99not_unsafe rate (%)66.7

Interactive version: theaggregate.ai/model?slug=gpt-5-6-luna-august-2026 · How It Works · Data refreshed daily, snapshot 2026-09-24.