GPT-5.6 Luna (Max): benchmark results
GPT-5.6 Luna evaluated at the max reasoning-effort setting. Provider: OpenAI. Released 2026-07-09. Access: API.
Unified ELO 1700 ± 1, rank #109 of 1761 rated models, from 92 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LiveBench Code Completion | 86.96 | Score | 99.1 |
| AA Long Context Reasoning | 83.67 | Accuracy (%) | 98.3 |
| Vals AI TaxEval v2 | 76.17 | Accuracy (%) | 96.2 |
| OTIS Mock AIME 2024-25 | 98.33 | Accuracy (%) | 95.3 |
| AA CritPt | 20.57 | Accuracy (%) | 93 |
| Artificial Analysis Intelligence Index | 43.44 | Intelligence Index | 93 |
| Vals AI IOI | 72.92 | Accuracy (%) | 93 |
| FrontierMath - Tiers 1-3 (v2) | 82.11 | Accuracy (%, 285 private v2 problems) | 92.9 |
| AA-LCR | 78.33 | Accuracy (self-reported) | 92 |
| AA GPQA Diamond | 91.11 | Accuracy (%) | 91.7 |
| Aikido CVE Rediscovery | 74.4 | Recall (%) | 91.3 |
| ZeroBench | 21 | Score (%) | 91.3 |
Interactive version: theaggregate.ai/model?slug=gpt-5-6-luna-max · How It Works · Data refreshed daily, snapshot 2026-09-05.