PRIME (Process-Outcome Alignment) - Biology: leaderboard

Metric: Overall accuracy (%) of the model as a verifier on PRIME's 397 biology items (textbook and exam problems with a model-written solution, kept where GPT-OSS-120B's eight verification verdicts disagreed): the model must accept a solution only if both its derivation and its final answer are correct, scored against expert process-outcome labels; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 33 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro91.44#145
2DeepSeek V3.2 (Thinking)90.43#198 (DeepSeek V3.2)
3Grok 489.42#169
4GPT-5.2 Instant89.42#205
5GPT-5.2 (Thinking)89.17#105 (GPT-5.2)
6Claude Sonnet 4.5 (Thinking)88.66#138 (Claude Sonnet 4.5)
7GPT-5 (Thinking)88.16#91 (GPT-5)
8Kimi K2 (Thinking)87.66#236 (Kimi K2)
9Qwen 3 30B A3B87.66#488
10Qwen 3 4B 2507 (Thinking)86.65#525 (Qwen 3 4B 2507)
11Claude Opus 4 (Thinking)86.65#155 (Claude Opus 4)
12Gemini 3 Pro85.39#77
13Claude Sonnet 4 (Thinking)85.39#194 (Claude Sonnet 4)
14GPT-OSS-20B85.14#499
15Claude Opus 4.5 (Thinking)85.14#79 (Claude Opus 4.5)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=prime-process-outcome-alignment-biology · How It Works · Data refreshed daily, snapshot 2026-10-11.