PRMBench — leaderboard

Process Reward Model benchmark evaluating fine-grained mathematical reasoning step verification across multiple error types.

Metric: Overall Score. Source: prmbench.github.io. Status: saturated. 26 models tracked.

Top models

#ModelScore
1O1 Mini68.8
2GPT-4o66.8
3Gemini 2.0 Flash (Preview)66
4DeepSeek R1 Distill Qwen 32B60.2
5DeepSeek R1 Distill Llama 70B57.5
6DeepSeek R1 Distill Llama 8B52.7
7DeepSeek-R1-Distill-Qwen-7B52.6

Interactive version: theaggregate.ai/benchmark?slug=prmbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.