PRIME (Process-Outcome Alignment) - Physics: leaderboard

Metric: Overall accuracy (%) of the model as a verifier on PRIME's 448 physics items (textbook and exam problems with a model-written solution, kept where GPT-OSS-120B's eight verification verdicts disagreed): the model must accept a solution only if both its derivation and its final answer are correct, scored against expert process-outcome labels; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 33 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 4.5 (Thinking)89.96#138 (Claude Sonnet 4.5)
2GPT-5.2 (Thinking)89.96#105 (GPT-5.2)
3GPT-5 (Thinking)89.96#91 (GPT-5)
4Gemini 2.5 Pro88.17#145
5GPT-5.2 Instant88.17#205
6Grok 487.72#169
7DeepSeek V3.2 (Thinking)87.72#198 (DeepSeek V3.2)
8Claude Opus 4 (Thinking)87.28#155 (Claude Opus 4)
9Kimi K2 (Thinking)87.05#236 (Kimi K2)
10Claude Opus 4.5 (Thinking)86.61#79 (Claude Opus 4.5)
11Claude Sonnet 4 (Thinking)85.94#194 (Claude Sonnet 4)
12Gemini 3 Pro85.49#77
13GLM-4.685.27#246
14GPT-OSS-20B83.26#499
15Qwen 3 30B A3B83.26#488

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=prime-process-outcome-alignment-physics · How It Works · Data refreshed daily, snapshot 2026-10-11.