MathDuels - Author Rating: leaderboard

Metric: Author rating (Elo-style points, open scale): mean rated difficulty of the model's own valid problems; 19 models each author 30 hardened problems (six mathematical domains) and solve every problem authored by the others; answers checked symbolically, ill-posed problems removed by a GPT-5.4-high verifier; Rasch model on a shared rating scale anchored at Grok-4.1-fast-high = 1500, k = 30; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 19 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview) (High)1624
2GPT-5.4 (High)1473
3GPT-5.2 (High)1427
4Gemini 3 Flash (High)1401
5GPT-5.4 (Low)1315
6Claude Opus 4.6 (High)1307
7Claude Sonnet 4.6 (High)1254
8GPT-5.4 Mini (High)1218
9Gemini 3.1 Pro (Preview) (Low)1208
10Kimi K2.51158
11Qwen 3.5 397B A17B1123
12MiniMax-M2.71122
13GLM-51095
14GPT-5.4 Mini (Low)1091
15DeepSeek V3.21072

Interactive version: theaggregate.ai/benchmark?slug=mathduels-author-rating · How It Works · Data refreshed daily, snapshot 2026-10-07.