GPT-5.6 Luna (xHigh) — benchmark results
GPT-5.6 Luna evaluated at the xhigh reasoning-effort setting. Provider: OpenAI. Released 2026-07-09. Access: API.
Unified ELO 1901 ± 30, rank #64 of 1776 rated models, from 43 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Omniscience - Software Engineering (SWE) - Kotlin | 70 | Accuracy (%) | 96.4 |
| AA CritPt | 20.57 | Accuracy (%) | 96.1 |
| Artificial Analysis Intelligence Index | 49.07 | Intelligence Index | 95.9 |
| AA Omniscience - Software Engineering (SWE) - TypeScript | 75.56 | Accuracy (%) | 95.3 |
| AA Humanity's Last Exam | 35.59 | Accuracy (%) | 93.5 |
| GDPval-AA | 1539 | Elo | 93.5 |
| AA GDPval | 1539.11 | ELO | 93.3 |
| AA Omniscience - Software Engineering (SWE) - Python | 71 | Accuracy (%) | 93.2 |
| AA Omniscience - Software Engineering (SWE) - Go | 66 | Accuracy (%) | 93.1 |
| AA Long Context Reasoning | 69.67 | Accuracy (%) | 92.7 |
| AA GPQA Diamond | 89.49 | Accuracy (%) | 92.5 |
| AA SciCode | 50 | Accuracy (%) | 92.5 |
Interactive version: theaggregate.ai/model?slug=gpt-5-6-luna-xhigh · How the rankings work · Data refreshed daily, snapshot 2026-07-22.