GPT-6 (xHigh): benchmark results
Provider: OpenAI. Released 2026-09-03. Access: API.
Unified ELO 1811 ± 1, rank #3 of 1761 rated models, from 25 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA GPQA Diamond | 96.26 | Accuracy (%) | 100 |
| DeepSWE | 74.1 | Pass@1 (%) | 100 |
| DeepsecBench | 37.79 | Recall-weighted F2 score (%) | 100 |
| Artificial Analysis Intelligence Index | 54.31 | Intelligence Index | 99.7 |
| AA CritPt | 31.43 | Accuracy (%) | 99.6 |
| AA Omniscience | 43.42 | Score | 99.6 |
| AA Omniscience - Software Engineering (SWE) | 91 | Accuracy (%) | 99.6 |
| ARC-AGI-2 | 93.33 | Accuracy (%) | 99.5 |
| ARC-AGI-1 | 98.5 | Accuracy (%) | 99.3 |
| AA MMMU-Pro | 86.24 | Accuracy (%) | 99.2 |
| AA Humanity's Last Exam | 54.59 | Accuracy (%) | 99 |
| AA Omniscience - Health | 53.7 | Accuracy (%) | 98.9 |
Interactive version: theaggregate.ai/model?slug=gpt-6-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-05.