PRIME (Process-Outcome Alignment) - Math: leaderboard

Metric: Overall accuracy (%) of the model as a verifier on PRIME's 900 mathematics items (textbook and exam problems with a model-written solution, kept where GPT-OSS-120B's eight verification verdicts disagreed): the model must accept a solution only if both its derivation and its final answer are correct, scored against expert process-outcome labels; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 33 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 (Thinking)86.67#91 (GPT-5)
2Gemini 2.5 Pro86.11#145
3GPT-5.2 (Thinking)86#105 (GPT-5.2)
4Grok 485.78#169
5DeepSeek V3.2 (Thinking)83.89#198 (DeepSeek V3.2)
6Kimi K2 (Thinking)82.56#236 (Kimi K2)
7GPT-5.2 Instant81.89#205
8Gemini 3 Pro81.67#77
9Claude Sonnet 4.5 (Thinking)81#138 (Claude Sonnet 4.5)
10GLM-4.681#246
11Claude Sonnet 4 (Thinking)81#194 (Claude Sonnet 4)
12Claude Opus 4 (Thinking)80.89#155 (Claude Opus 4)
13Qwen 3 30B A3B79.67#488
14Claude Opus 4.5 (Thinking)79.33#79 (Claude Opus 4.5)
15Qwen 3 4B 2507 (Thinking)76.22#525 (Qwen 3 4B 2507)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=prime-process-outcome-alignment-math · How It Works · Data refreshed daily, snapshot 2026-10-11.