MedPRMBench (Generative): leaderboard
Metric: PRMScore (%): mean of the F1 for erroneous steps and the F1 for correct steps (weights 0.5 and 0.5) in step-level error detection over the erroneous reasoning chains of the 6,500 MedPRMBench test instances (14 injected medical reasoning error types), generative protocol: the API model receives the medical question with numbered reasoning steps and outputs one plus or minus symbol per step; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 75.4 |
| 2 | GPT-5.2 | 74.6 |
| 3 | Claude Opus 4.5 | 73.7 |
| 4 | GLM-4.7 | 72.4 |
| 5 | DeepSeek V3.1 | 70.7 |
| 6 | Qwen 3 Max | 70.4 |
| 7 | DeepSeek R1 | 67.8 |
| 8 | DeepSeek V3.2 | 65.4 |
| 9 | Gemini 3.1 Pro (Preview) | 63.6 |
Interactive version: theaggregate.ai/benchmark?slug=medprmbench-generative · How It Works · Data refreshed daily, snapshot 2026-10-07.