MathDuels: leaderboard
Metric: Composite rating (Elo-style points, open scale): the mean of the solver and author ratings; 19 models each author 30 hardened problems (six mathematical domains) and solve every problem authored by the others; answers checked symbolically, ill-posed problems removed by a GPT-5.4-high verifier; Rasch model on a shared rating scale anchored at Grok-4.1-fast-high = 1500, k = 30; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 19 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) (High) | 1919 |
| 2 | GPT-5.4 (High) | 1870 |
| 3 | GPT-5.2 (High) | 1737 |
| 4 | Claude Opus 4.6 (High) | 1675 |
| 5 | Gemini 3 Flash (High) | 1673 |
| 6 | GPT-5.4 (Low) | 1637 |
| 7 | Claude Sonnet 4.6 (High) | 1593 |
| 8 | GPT-5.4 Mini (High) | 1570 |
| 9 | Qwen 3.5 397B A17B | 1548 |
| 10 | Kimi K2.5 | 1542 |
| 11 | Gemini 3.1 Pro (Preview) (Low) | 1503 |
| 12 | Grok 4.20 (High) | 1485 |
| 13 | DeepSeek V3.2 | 1449 |
| 14 | GLM-5 | 1438 |
| 15 | MiniMax-M2.7 | 1372 |
Interactive version: theaggregate.ai/benchmark?slug=mathduels · How It Works · Data refreshed daily, snapshot 2026-10-07.