GPT-5.6 Luna (Max) — benchmark results

GPT-5.6 Luna evaluated at the max reasoning-effort setting. Provider: OpenAI. Released 2026-07-09. Access: API.

Unified ELO 1984 ± 27, rank #29 of 1776 rated models, from 49 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA Long Context Reasoning74Accuracy (%)98.3
Artificial Analysis Intelligence Index51.24Intelligence Index97.2
AA CritPt20.57Accuracy (%)96.1
AA Omniscience - Software Engineering (SWE) - TypeScript77.78Accuracy (%)96
AA GPQA Diamond91.11Accuracy (%)95.9
GDPval-AA1584Elo95.9
OTIS Mock AIME 2024-2598.33Accuracy (%)95.8
AA GDPval1584ELO95.7
AA SciCode52.55Accuracy (%)95.3
AA Humanity's Last Exam37.21Accuracy (%)95.1
AA Omniscience - Software Engineering (SWE) - Kotlin66Accuracy (%)94.6
AA Omniscience - Software Engineering (SWE) - JavaScript73.64Accuracy (%)93.8

Interactive version: theaggregate.ai/model?slug=gpt-5-6-luna-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.