MedPRMBench (Generative): leaderboard

Metric: PRMScore (%): mean of the F1 for erroneous steps and the F1 for correct steps (weights 0.5 and 0.5) in step-level error detection over the erroneous reasoning chains of the 6,500 MedPRMBench test instances (14 injected medical reasoning error types), generative protocol: the API model receives the medical question with numbered reasoning steps and outputs one plus or minus symbol per step; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 9 models tracked.

Top models

#ModelScore
1GPT-5.475.4
2GPT-5.274.6
3Claude Opus 4.573.7
4GLM-4.772.4
5DeepSeek V3.170.7
6Qwen 3 Max70.4
7DeepSeek R167.8
8DeepSeek V3.265.4
9Gemini 3.1 Pro (Preview)63.6

Interactive version: theaggregate.ai/benchmark?slug=medprmbench-generative · How It Works · Data refreshed daily, snapshot 2026-10-07.