ChatGPT: benchmark results
Provider: OpenAI. Access: API.
Unified ELO 1514 ± 18, rank #1010 of 2656 rated models, from 23 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| InstructEval - Alignment (HHH) | 86.6 | Average (%, harmless/helpful/honest) | 100 |
| InstructEval - Problem Solving | 64.5 | Average (%, MMLU/BBH/DROP/CRASS/HumanEval) | 100 |
| EvoEval Combine | 33 | Pass@1 (%) | 94 |
| EvoEval Subtle | 70 | Pass@1 (%) | 94 |
| MoralChoice | 100 | MoralChoice Acc (%) | 92.5 |
| EvoEval | 57.23 | Pass@1 (%) | 90 |
| EvoEval Tool Use | 64 | Pass@1 (%) | 89 |
| EvoEval Creative | 42 | Pass@1 (%) | 86 |
| ShoppingMMLU | 65.3 | Average Score (%) | 84.6 |
| EvoEval Verbose | 70.73 | Pass@1 (%) | 84 |
| EvoEval Concise | 68.29 | Pass@1 (%) | 78 |
| EvoEval Difficult | 33 | Pass@1 (%) | 74 |
Interactive version: theaggregate.ai/model?slug=chatgpt · How It Works · Data refreshed daily, snapshot 2026-09-19.