OpenRouter GPQA Diamond: leaderboard

OpenRouter's own GPQA Diamond run: 198 graduate-level science questions executed by its native benchmark harness against the production endpoints it serves, so the score reflects the deployed model rather than a vendor-reported figure.

Metric: Accuracy (%). Source: openrouter.ai. Status: saturation imminent. 134 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)94.4
2GPT-694.4
3Gemini 3.7 Flash94.3
4GPT-5.593.8
5GPT-5.6 Pro Sol93.8
6Grok 4.693.3
7Gemini 3.5 Flash92.8
8Gemini 3.6 Flash92.8
9GPT-5.6 Sol91.9
10Kimi K391.9
11MiniMax-M390.6
12GPT-5.490.3
13GPT-5.6 Pro Luna90.3
14Claude Fable 5.190.1
15Claude Opus 4.889.7

Interactive version: theaggregate.ai/benchmark?slug=openrouter-gpqa-diamond · How It Works · Data refreshed daily, snapshot 2026-09-20.