Open LLM Leaderboard - GPQA: leaderboard
Metric: Score. Source: huggingface.co. 4576 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | 70B-L3.3-Cirrus-x1 | 26.62 |
| 2 | Mistral Large 2 (Nov) Instruct (2411) | 24.94 |
| 3 | Qwen 2.5 32B | 21.59 |
| 4 | Phi-4 | 20.81 |
| 5 | Qwen 2.5 72B | 20.69 |
| 6 | phi-4-14B | 20.47 |
| 7 | calme-3.2-instruct-78B | 20.36 |
| 8 | Llama 3 70B | 19.69 |
| 9 | Saka-14B | 19.46 |
| 10 | Qwen 2 72B | 19.24 |
| 11 | Llama-3.1-Nemotron-lorablated-70B | 18.79 |
| 12 | Llama 3.1 70B | 18.34 |
| 13 | DeepSeek R1 Distill Qwen 14B | 18.34 |
| 14 | Qwen 2 VL 72B Instruct | 18.34 |
| 15 | Mistral-Small-24B-Base-2501 | 18.34 |
Interactive version: theaggregate.ai/benchmark?slug=open-llm-leaderboard-gpqa · How It Works · Data refreshed daily, snapshot 2026-09-05.