GPT-6 Luna (Max): benchmark results
Provider: OpenAI. Access: API.
Unified ELO 1707 ± 1, rank #141 of 3363 rated models, from 89 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Long Context Reasoning | 83.33 | Accuracy (%) | 96.7 |
| MLCR-AA - Accuracy | 93.75 | Accuracy (%) | 96.6 |
| Artificial Analysis Intelligence Index | 37.26 | Intelligence Index | 90.9 |
| AA CritPt | 19.43 | Accuracy (%) | 90.8 |
| AA-Omniscience Index - Software Engineering (SWE) - JavaScript | 56.36 | Omniscience Index | 90.7 |
| AA-Omniscience Index - Software Engineering (SWE) - Kotlin | 40 | Omniscience Index | 88.8 |
| AA-Omniscience Index - Software Engineering (SWE) - Swift | 52 | Omniscience Index | 88.8 |
| AA-Omniscience Index - Software Engineering (SWE) - Python | 50.5 | Omniscience Index | 88.5 |
| AA Humanity's Last Exam | 38.51 | Accuracy (%) | 87.6 |
| LiveBench Connections | 100 | Score | 87.1 |
| AA-Omniscience Index - Software Engineering (SWE) - PHP | 48 | Omniscience Index | 86.7 |
| AA-Omniscience Index - Software Engineering (SWE) - TypeScript | 46.67 | Omniscience Index | 86.2 |
Interactive version: theaggregate.ai/model?slug=gpt-6-luna-max · How It Works · Data refreshed daily, snapshot 2026-09-23.