GPT-5.6 Luna (Max): benchmark results

GPT-5.6 Luna evaluated at the max reasoning-effort setting. Provider: OpenAI. Released 2026-07-09. Access: API.

Unified ELO 1700 ± 1, rank #109 of 1761 rated models, from 92 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LiveBench Code Completion86.96Score99.1
AA Long Context Reasoning83.67Accuracy (%)98.3
Vals AI TaxEval v276.17Accuracy (%)96.2
OTIS Mock AIME 2024-2598.33Accuracy (%)95.3
AA CritPt20.57Accuracy (%)93
Artificial Analysis Intelligence Index43.44Intelligence Index93
Vals AI IOI72.92Accuracy (%)93
FrontierMath - Tiers 1-3 (v2)82.11Accuracy (%, 285 private v2 problems)92.9
AA-LCR78.33Accuracy (self-reported)92
AA GPQA Diamond91.11Accuracy (%)91.7
Aikido CVE Rediscovery74.4Recall (%)91.3
ZeroBench21Score (%)91.3

Interactive version: theaggregate.ai/model?slug=gpt-5-6-luna-max · How It Works · Data refreshed daily, snapshot 2026-09-05.