GDPval (OpenAI Evals) — leaderboard

220 real-world tasks from 44 occupations across the top 9 GDP sectors. Built with industry professionals averaging 14 years of experience. Evaluates correctness, structure, style, and relevance.

Metric: Win Rate (%). Source: evals.openai.com. Status: saturation imminent. 16 models tracked.

Top models

#ModelScore
1GPT-5.2 (High)49.88
2Claude Opus 4.545.37
3Claude Opus 4.143.77
4Claude Sonnet 4.542.63
5Gemini 3 Pro40.25
6GPT-5 (High)34.77
7O3 (High)30.75
8O4 Mini (High)25.15
9Gemini 2.5 Pro23.42
10Grok 421.23
11GPT-5 (Medium)14.96
12GPT-5 (Low)14.02
13O3 (Medium)12.53
14O3 (Low)11.8
15GPT-4o9.85

Interactive version: theaggregate.ai/benchmark?slug=gdpval-openai-evals · How the rankings work · Data refreshed daily, snapshot 2026-07-22.