GPT-4 (0613): benchmark results
June 13, 2023 GPT-4 API snapshot, kept separate from other GPT-4 rows when sources report it explicitly. Provider: OpenAI. Released 2023-06-13. Access: API.
Unified ELO 1537 ± 1, rank #501 of 1392 rated models, from 80 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CommonGen | 45.11 | Overall (%) | 100 |
| FastEval | 77.78 | Total Score | 100 |
| HELM v2 Lite - HumanEval (Code) | 73 | Pass@1 (%) | 100 |
| HELM v2 Lite - LegalBench | 71.01 | Exact Match (%) | 100 |
| HELM v2 Lite - MedQA | 88.46 | Exact Match (%) | 100 |
| InfiBench | 70.64 | Score (%) | 100 |
| LogicKor - Reasoning | 9.42 | Score (0-10) | 98.3 |
| MAGI-Hard | 77.85 | Accuracy (%, 2024-05 snapshot) | 98.3 |
| HELM Lite | 90.89 | Mean win rate (self-reported) | 97.4 |
| JustEval - Clarity | 4.99 | Score (1-5) | 96.7 |
| JustEval - Factuality | 4.9 | Score (1-5) | 96.7 |
| LogicKor - Grammar | 8.35 | Score (0-10) | 96.6 |
Interactive version: theaggregate.ai/model?slug=gpt-4-0613 · How It Works · Data refreshed daily, snapshot 2026-09-05.