Open-R1 Eval Leaderboard — leaderboard

Open-R1 released-model reasoning leaderboard aggregating LightEval results across math, GPQA, AIME, MATH-500, and related reasoning tasks.

Metric: Average Accuracy (%). Source: huggingface.co. Status: saturation imminent. 37 models tracked.

Top models

#ModelScore
1Qwen 3 32B73.74
2Qwen 3 30B A3B72.94
3Qwen 3 14B72.67
4QwQ-32B70.58
5Qwen 3 8B69.79
6DeepSeek R1 Distill Qwen 32B65.96
7DeepSeek R1 Distill Llama 70B65.75
8Qwen 3 4B65.61
9DeepSeek R1 Distill Qwen 14B63.12
10DeepSeek-R1-Distill-Qwen-7B51.57
11QwQ 32B-Preview51.01
12Qwen 3 1.7B48.95
13DeepSeek R1 Distill Llama 8B47.77
14OpenThinker-7B37.37
15DeepSeek R1 Distill Qwen 1.5B34.46

Interactive version: theaggregate.ai/benchmark?slug=open-r1-eval-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.