GPT-5.6 Luna (xHigh): benchmark results
GPT-5.6 Luna evaluated at the xhigh reasoning-effort setting. Provider: OpenAI. Released 2026-07-09. Access: API.
Unified ELO 1692 ± 1, rank #125 of 1761 rated models, from 76 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Chatbot Arena (Text - German) | 1486 | Arena Score | 94.2 |
| AA Long Context Reasoning | 81.67 | Accuracy (%) | 94 |
| Chatbot Arena (Text - Math) | 1478 | Arena Score | 93.5 |
| LLM Chess (Saplin) | 1185.3 | ELO | 93.2 |
| AA CritPt | 20.57 | Accuracy (%) | 93 |
| Artificial Analysis Intelligence Index | 41.58 | Intelligence Index | 91.2 |
| Chess Bench LLM | 1469 | Lichess Rating | 90.6 |
| Chatbot Arena (Text - Japanese) | 1449 | Arena Score | 89.4 |
| AA Humanity's Last Exam | 36.98 | Accuracy (%) | 87.9 |
| AA Omniscience - Software Engineering (SWE) | 65.9 | Accuracy (%) | 87.7 |
| AA GPQA Diamond | 89.49 | Accuracy (%) | 87.6 |
| AA Omniscience - Science, Engineering & Mathematics | 45 | Accuracy (%) | 86.7 |
Interactive version: theaggregate.ai/model?slug=gpt-5-6-luna-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-05.