GPT-5.6 Luna (High): benchmark results

GPT-5.6 Luna evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2026-07-09. Access: API.

Unified ELO 1681 ± 1, rank #148 of 1761 rated models, from 30 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
PM-LLM-Benchmark37.2Score93.4
AA Long Context Reasoning80.33Accuracy (%)90.8
AA CritPt16.57Accuracy (%)89.7
Artificial Analysis Intelligence Index37.36Intelligence Index89
AA Omniscience - Software Engineering (SWE)67.2Accuracy (%)88.9
AA GPQA Diamond89.19Accuracy (%)86.7
LLM Chess (Saplin)919.4ELO86.3
AA Omniscience - Science, Engineering & Mathematics44.3Accuracy (%)85.2
AA-Omniscience Accuracy41.78Accuracy (%)85.2
AA Humanity's Last Exam33.41Accuracy (%)84.7
AA Omniscience - Business33.9Accuracy (%)84.2
AA Omniscience - Humanities & Social Sciences39.2Accuracy (%)83.9

Interactive version: theaggregate.ai/model?slug=gpt-5-6-luna-high · How It Works · Data refreshed daily, snapshot 2026-09-05.