Prim - Discovery: leaderboard

Metric: Discovery score (%): how well the model states the problem's mathematical primitive (the key organizing idea) from the problem alone, judged against the human-verified gold primitive by GPT-5.4-High on a graded rubric, over 182 problems; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScore
1GPT-5.4 (High)82.42
2GPT-5.4 Mini (High)59.34
3GPT-5.4 Nano (High)41.76
4GPT-OSS-20B (High)34.62
5Qwen 3.5 27B28.57
6Qwen 3.6 27B24.73
7Qwen 3.5 9B13.19
8DeepSeek-R1-Distill-Qwen-7B7.14
9Qwen 3.5 4B6.59
10DeepSeek R1 0528 Qwen3 8B6.59
11DeepSeek R1 Distill Qwen 32B6.04
12DeepSeek R1 Distill Qwen 14B4.4

Interactive version: theaggregate.ai/benchmark?slug=prim-discovery · How It Works · Data refreshed daily, snapshot 2026-10-04.