OpenRouter GPQA Diamond — leaderboard

OpenRouter's own GPQA Diamond run: 198 graduate-level science questions executed by its native benchmark harness against the production endpoints it serves, so the score reflects the deployed model rather than a vendor-reported figure.

Metric: Accuracy (%). Source: openrouter.ai. Status: saturation imminent. 105 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)94.3
2GPT-5.6 Pro Sol93.8
3GPT-5.593.3
4Kimi K393.3
5Gemini 3.5 Flash93.1
6Gemini 3.6 Flash92.4
7GPT-5.6 Sol91.4
8MiniMax-M391.2
9GPT-5.490.4
10Claude Opus 4.889.8
11Nova Micro89.7
12GPT-5.6 Terra89.6
13Claude Opus 589.2
14Claude Opus 4.788.5
15GPT-5.6 Luna88.4

Interactive version: theaggregate.ai/benchmark?slug=openrouter-gpqa-diamond · How It Works · Data refreshed daily, snapshot 2026-08-06.