PRIME (Process-Outcome Alignment) - Chemistry: leaderboard

Metric: Overall accuracy (%) of the model as a verifier on PRIME's 782 chemistry items (textbook and exam problems with a model-written solution, kept where GPT-OSS-120B's eight verification verdicts disagreed): the model must accept a solution only if both its derivation and its final answer are correct, scored against expert process-outcome labels; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 32 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro90.15#145
2Grok 489#169
3Kimi K2 (Thinking)89#236 (Kimi K2)
4GPT-5 (Thinking)88.62#91 (GPT-5)
5DeepSeek V3.2 (Thinking)88.49#198 (DeepSeek V3.2)
6GPT-5.2 (Thinking)88.49#105 (GPT-5.2)
7Claude Sonnet 4.5 (Thinking)88.11#138 (Claude Sonnet 4.5)
8Claude Opus 4 (Thinking)86.57#155 (Claude Opus 4)
9GPT-5.2 Instant86.45#205
10Gemini 3 Pro86.19#77
11Claude Sonnet 4 (Thinking)85.29#194 (Claude Sonnet 4)
12GLM-4.684.91#246
13Qwen 3 4B 2507 (Thinking)83.76#525 (Qwen 3 4B 2507)
14Qwen 3 30B A3B83.38#488
15Claude Opus 4.5 (Thinking)81.84#79 (Claude Opus 4.5)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=prime-process-outcome-alignment-chemistry · How It Works · Data refreshed daily, snapshot 2026-10-11.