PRMBench: leaderboard

Process Reward Model benchmark evaluating fine-grained mathematical reasoning step verification across multiple error types.

Metric: Overall Score. Source: prmbench.github.io. Status: saturated. 26 models tracked.

Top models

#ModelScore
1O1 Mini68.8
2GPT-4o66.8
3DeepSeek R1 Distill Qwen 32B60.2
4DeepSeek R1 Distill Llama 70B57.5
5Qwen2.5-Math-72B57.4
6DeepSeek R1 Distill Llama 8B52.7
7DeepSeek-R1-Distill-Qwen-7B52.6

Interactive version: theaggregate.ai/benchmark?slug=prmbench · How It Works · Data refreshed daily, snapshot 2026-09-05.