GDPval (OpenAI Evals) — leaderboard
220 real-world tasks from 44 occupations across the top 9 GDP sectors. Built with industry professionals averaging 14 years of experience. Evaluates correctness, structure, style, and relevance.
Metric: Win Rate (%). Source: evals.openai.com. Status: saturation imminent. 16 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 (High) | 49.88 |
| 2 | Claude Opus 4.5 | 45.37 |
| 3 | Claude Opus 4.1 | 43.77 |
| 4 | Claude Sonnet 4.5 | 42.63 |
| 5 | Gemini 3 Pro | 40.25 |
| 6 | GPT-5 (High) | 34.77 |
| 7 | O3 (High) | 30.75 |
| 8 | O4 Mini (High) | 25.15 |
| 9 | Gemini 2.5 Pro | 23.42 |
| 10 | Grok 4 | 21.23 |
| 11 | GPT-5 (Medium) | 14.96 |
| 12 | GPT-5 (Low) | 14.02 |
| 13 | O3 (Medium) | 12.53 |
| 14 | O3 (Low) | 11.8 |
| 15 | GPT-4o | 9.85 |
Interactive version: theaggregate.ai/benchmark?slug=gdpval-openai-evals · How the rankings work · Data refreshed daily, snapshot 2026-07-22.