MathDuels - Author Rating: leaderboard
Metric: Author rating (Elo-style points, open scale): mean rated difficulty of the model's own valid problems; 19 models each author 30 hardened problems (six mathematical domains) and solve every problem authored by the others; answers checked symbolically, ill-posed problems removed by a GPT-5.4-high verifier; Rasch model on a shared rating scale anchored at Grok-4.1-fast-high = 1500, k = 30; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 19 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) (High) | 1624 |
| 2 | GPT-5.4 (High) | 1473 |
| 3 | GPT-5.2 (High) | 1427 |
| 4 | Gemini 3 Flash (High) | 1401 |
| 5 | GPT-5.4 (Low) | 1315 |
| 6 | Claude Opus 4.6 (High) | 1307 |
| 7 | Claude Sonnet 4.6 (High) | 1254 |
| 8 | GPT-5.4 Mini (High) | 1218 |
| 9 | Gemini 3.1 Pro (Preview) (Low) | 1208 |
| 10 | Kimi K2.5 | 1158 |
| 11 | Qwen 3.5 397B A17B | 1123 |
| 12 | MiniMax-M2.7 | 1122 |
| 13 | GLM-5 | 1095 |
| 14 | GPT-5.4 Mini (Low) | 1091 |
| 15 | DeepSeek V3.2 | 1072 |
Interactive version: theaggregate.ai/benchmark?slug=mathduels-author-rating · How It Works · Data refreshed daily, snapshot 2026-10-07.