GPT-6 (Medium): benchmark results
Provider: OpenAI. Released 2026-09-03. Access: API.
Unified ELO 1778 ± 1, rank #13 of 1761 rated models, from 24 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| FrontierMath - Tier 4 (v2) | 97.56 | Accuracy (%, 41 private v2 problems) | 100 |
| AA Omniscience - Software Engineering (SWE) | 90.2 | Accuracy (%) | 99 |
| AA Omniscience | 42.22 | Score | 98.8 |
| Artificial Analysis Intelligence Index | 52.25 | Intelligence Index | 98.6 |
| AA GPQA Diamond | 93.94 | Accuracy (%) | 98.5 |
| ARC-AGI-2 | 92.08 | Accuracy (%) | 98.4 |
| AA Humanity's Last Exam | 52.73 | Accuracy (%) | 98.2 |
| AA-Omniscience Accuracy | 60.57 | Accuracy (%) | 98.2 |
| AA CritPt | 29.14 | Accuracy (%) | 98.1 |
| AA Omniscience - Health | 52.4 | Accuracy (%) | 98 |
| AA MMMU-Pro | 85.09 | Accuracy (%) | 97.7 |
| AA Omniscience - Business | 50 | Accuracy (%) | 97.7 |
Interactive version: theaggregate.ai/model?slug=gpt-6-medium · How It Works · Data refreshed daily, snapshot 2026-09-05.