HELM Capabilities - GPQA — leaderboard

Metric: COT correct. Source: crfm.stanford.edu. 51 models tracked.

Top models

#ModelScore
1O3 (2025-04-16)75.34
2Gemini 2.5 Pro (Preview 03-25)74.89
3O4 Mini (2025-04-16)73.54
4Grok 472.65
5Grok 3 Mini Beta67.49
6DeepSeek R1 052866.59
7Claude Opus 4 (20250514)66.59
8Palmyra X566.14
9GPT-4.1 (2025-04-14)65.92
10Kimi K265.25
11Grok 365.02
12Llama 4 Maverick Instruct FP865.02
13Claude Sonnet 4 (20250514)64.35
14Qwen3 235B A22B FP8 Throughput62.33
15GPT-4.1 Mini61.43

Interactive version: theaggregate.ai/benchmark?slug=helm-capabilities-gpqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.