GPT-4 — benchmark results
OpenAI's GPT-4 flagship model (March 2023). Provider: OpenAI. Released 2023-03-14. Access: API.
Unified ELO 1580 ± 19, rank #531 of 1776 rated models, from 160 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AgentBoard | 70 | Progress Rate (self-reported) | 100 |
| BELLS | 91.7 | BELLS Score (%) | 100 |
| ClassEval | 37.6 | Class-Level Pass@1 (%) | 100 |
| CodeScope | 39.2 | Score | 100 |
| CodeScope - Code Generation | 31.47 | Score | 100 |
| CodeScope - Code Repair | 30.03 | Score | 100 |
| CodeScope - Code Summarization | 33.66 | Score | 100 |
| CodeScope - Code Translation | 31.29 | Score | 100 |
| CodeScope - Program Synthesis | 36.36 | Score | 100 |
| CyberBench (NLP) | 74.03 | Avg Score (%) | 100 |
| EvoEval Combine | 53 | Pass@1 (%) | 100 |
| EvoEval Concise | 81.1 | Pass@1 (%) | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-4 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.