GPT-5.6 Luna (Max) — benchmark results
GPT-5.6 Luna evaluated at the max reasoning-effort setting. Provider: OpenAI. Released 2026-07-09. Access: API.
Unified ELO 1984 ± 27, rank #29 of 1776 rated models, from 49 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Long Context Reasoning | 74 | Accuracy (%) | 98.3 |
| Artificial Analysis Intelligence Index | 51.24 | Intelligence Index | 97.2 |
| AA CritPt | 20.57 | Accuracy (%) | 96.1 |
| AA Omniscience - Software Engineering (SWE) - TypeScript | 77.78 | Accuracy (%) | 96 |
| AA GPQA Diamond | 91.11 | Accuracy (%) | 95.9 |
| GDPval-AA | 1584 | Elo | 95.9 |
| OTIS Mock AIME 2024-25 | 98.33 | Accuracy (%) | 95.8 |
| AA GDPval | 1584 | ELO | 95.7 |
| AA SciCode | 52.55 | Accuracy (%) | 95.3 |
| AA Humanity's Last Exam | 37.21 | Accuracy (%) | 95.1 |
| AA Omniscience - Software Engineering (SWE) - Kotlin | 66 | Accuracy (%) | 94.6 |
| AA Omniscience - Software Engineering (SWE) - JavaScript | 73.64 | Accuracy (%) | 93.8 |
Interactive version: theaggregate.ai/model?slug=gpt-5-6-luna-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.