SuperGPQA — leaderboard

Graduate-level QA benchmark spanning 285 disciplines. 10,000+ questions at graduate-level difficulty from verified experts. Tests deep, specialized knowledge beyond typical benchmarks.

Metric: Accuracy (%). Source: supergpqa.github.io. Status: saturation imminent. 69 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)73.75
2GPT-5.2 Pro67.13
3GPT-5.1 (Thinking)66.35
4GPT-565.63
5Gemini 2.5 Pro64.51
6Grok 462.99
7DeepSeek R161.82
8Claude Opus 4 (20250514)61.43
9O1 (2024-12-17)60.24
10DeepSeek V3.159.32
11Gemini 2.5 Flash58.91
12GPT-5 Mini58.82
13Kimi K2 (0711)58.08
14GPT-5 Chat57.36
15Claude Sonnet 4 (20250514)57.01

Interactive version: theaggregate.ai/benchmark?slug=supergpqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.