Prim - Execution: leaderboard

Metric: Execution accuracy (%): exact-match accuracy on the 182 problems when the gold mathematical primitive is given with the problem; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScore
1GPT-5.4 (High)86.81
2Qwen 3.6 27B78.57
3Qwen 3.5 27B71.98
4GPT-5.4 Mini (High)71.98
5GPT-5.4 Nano (High)68.68
6Qwen 3.5 9B61.54
7GPT-OSS-20B (High)60.99
8Qwen 3.5 4B56.04
9DeepSeek R1 Distill Qwen 32B43.41
10DeepSeek R1 0528 Qwen3 8B39.56
11DeepSeek R1 Distill Qwen 14B39.01
12DeepSeek-R1-Distill-Qwen-7B32.42

Interactive version: theaggregate.ai/benchmark?slug=prim-execution · How It Works · Data refreshed daily, snapshot 2026-10-04.