MathDuels - Solver Rating: leaderboard

Metric: Solver rating (Elo-style points, open scale): Rasch ability from solving the other models' problems; 19 models each author 30 hardened problems (six mathematical domains) and solve every problem authored by the others; answers checked symbolically, ill-posed problems removed by a GPT-5.4-high verifier; Rasch model on a shared rating scale anchored at Grok-4.1-fast-high = 1500, k = 30; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 19 models tracked.

Top models

#ModelScore
1GPT-5.4 (High)2268
2Gemini 3.1 Pro (Preview) (High)2214
3GPT-5.2 (High)2047
4Claude Opus 4.6 (High)2043
5Qwen 3.5 397B A17B1972
6GPT-5.4 (Low)1958
7Grok 4.20 (High)1950
8Gemini 3 Flash (High)1944
9Claude Sonnet 4.6 (High)1932
10Kimi K2.51925
11GPT-5.4 Mini (High)1922
12DeepSeek V3.21826
13Gemini 3.1 Pro (Preview) (Low)1797
14GLM-51781
15Step 3.5 Flash1624

Interactive version: theaggregate.ai/benchmark?slug=mathduels-solver-rating · How It Works · Data refreshed daily, snapshot 2026-10-07.