AA GPQA Diamond — leaderboard

Artificial Analysis independent evaluation of GPQA Diamond: 198 hardest graduate-level science questions where PhD experts achieve only 65%.

Metric: Accuracy (%). Source: artificialanalysis.ai. Status: saturated. 544 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)94.14
2GPT-5.6 Sol (Max)94.14
3GPT-5.5 (xHigh)93.54
4Kimi K393.54
5GPT-5.5 (High)93.23
6GPT-5.6 Sol (xHigh)93.13
7Grok 4.5 (High)93.13
8MiniMax-M392.93
9GPT-5.6 Sol (High)92.83
10GPT-5.5 (Medium)92.63
11GPT-5.6 Sol (Medium)92.63
12GPT-5.6 Terra (Max)92.53
13Qwen 3.7 Max92.32
14Gemini 3.5 Flash (High)92.22
15Gemini 3.5 Flash (Medium)92.12

Interactive version: theaggregate.ai/benchmark?slug=aa-gpqa-diamond · How the rankings work · Data refreshed daily, snapshot 2026-07-22.