MathDuels - Solver Rating: leaderboard
Metric: Solver rating (Elo-style points, open scale): Rasch ability from solving the other models' problems; 19 models each author 30 hardened problems (six mathematical domains) and solve every problem authored by the others; answers checked symbolically, ill-posed problems removed by a GPT-5.4-high verifier; Rasch model on a shared rating scale anchored at Grok-4.1-fast-high = 1500, k = 30; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 19 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 (High) | 2268 |
| 2 | Gemini 3.1 Pro (Preview) (High) | 2214 |
| 3 | GPT-5.2 (High) | 2047 |
| 4 | Claude Opus 4.6 (High) | 2043 |
| 5 | Qwen 3.5 397B A17B | 1972 |
| 6 | GPT-5.4 (Low) | 1958 |
| 7 | Grok 4.20 (High) | 1950 |
| 8 | Gemini 3 Flash (High) | 1944 |
| 9 | Claude Sonnet 4.6 (High) | 1932 |
| 10 | Kimi K2.5 | 1925 |
| 11 | GPT-5.4 Mini (High) | 1922 |
| 12 | DeepSeek V3.2 | 1826 |
| 13 | Gemini 3.1 Pro (Preview) (Low) | 1797 |
| 14 | GLM-5 | 1781 |
| 15 | Step 3.5 Flash | 1624 |
Interactive version: theaggregate.ai/benchmark?slug=mathduels-solver-rating · How It Works · Data refreshed daily, snapshot 2026-10-07.