Prim - Generation: leaderboard

Metric: Generation accuracy (%): exact match of the model's own zero-shot chain-of-thought final answer on 182 HLE-Verified mathematics problems; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScore
1GPT-5.4 (High)67.03
2Qwen 3.6 27B52.75
3GPT-5.4 Mini (High)50.55
4Qwen 3.5 27B47.8
5GPT-5.4 Nano (High)44.51
6GPT-OSS-20B (High)43.41
7Qwen 3.5 9B38.46
8Qwen 3.5 4B26.37
9DeepSeek R1 Distill Qwen 32B18.13
10DeepSeek R1 0528 Qwen3 8B17.03
11DeepSeek R1 Distill Qwen 14B15.93
12DeepSeek-R1-Distill-Qwen-7B13.19

Interactive version: theaggregate.ai/benchmark?slug=prim-generation · How It Works · Data refreshed daily, snapshot 2026-10-04.