MathDuels: leaderboard

Metric: Composite rating (Elo-style points, open scale): the mean of the solver and author ratings; 19 models each author 30 hardened problems (six mathematical domains) and solve every problem authored by the others; answers checked symbolically, ill-posed problems removed by a GPT-5.4-high verifier; Rasch model on a shared rating scale anchored at Grok-4.1-fast-high = 1500, k = 30; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 19 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview) (High)1919
2GPT-5.4 (High)1870
3GPT-5.2 (High)1737
4Claude Opus 4.6 (High)1675
5Gemini 3 Flash (High)1673
6GPT-5.4 (Low)1637
7Claude Sonnet 4.6 (High)1593
8GPT-5.4 Mini (High)1570
9Qwen 3.5 397B A17B1548
10Kimi K2.51542
11Gemini 3.1 Pro (Preview) (Low)1503
12Grok 4.20 (High)1485
13DeepSeek V3.21449
14GLM-51438
15MiniMax-M2.71372

Interactive version: theaggregate.ai/benchmark?slug=mathduels · How It Works · Data refreshed daily, snapshot 2026-10-07.