GPT-6: benchmark results
Provider: OpenAI. Released 2026-09-03. Access: API.
Unified ELO 1805 ± 1, rank #4 of 1392 rated models, from 30 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Benchmarks.bio - TxBench-AD | 56.9 | Pass Rate (%) | 100 |
| LLM Stats (ARC-AGI-3) | 99.9 | Score (%) | 100 |
| LLM Stats (Agents' Last Exam) | 59.3 | Score (%) | 100 |
| LLM Stats (BrowseComp) | 91.5 | Score (%) | 100 |
| LLM Stats (ExploitBench) | 100 | Score (%) | 100 |
| LLM Stats (ExploitGym) | 42.4 | Score (%) | 100 |
| LLM Stats (ScreenSpot Pro) | 92.7 | Score (%) | 100 |
| LLM Stats (Terminal-Bench 4.0) | 57.7 | Score (%) | 100 |
| LLM Stats Score | 60.72 | LLM Stats Score (conservative rating) | 100 |
| RuneBench | 34063 | Total Peak XP Rate (XP/min) | 100 |
| Terminal-Bench 4.0 | 58.2 | Accuracy (%) | 100 |
| Vellum - ARC-AGI-2 | 95 | Score (%) | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-6 · How It Works · Data refreshed daily, snapshot 2026-09-05.