LiveMathematicianBench — leaderboard

Continuously updated benchmark testing LLMs on understanding cutting-edge mathematical theorems from newly published arXiv preprints. Multiple-choice questions constructed from theorem statements and proof sketches.

Metric: Accuracy (%). Source: livemathematicianbench.github.io. Status: saturation imminent. 7 models tracked.

Top models

#ModelScore
1GPT-5.458.6
2Gemini 3.1 Pro (Preview)54.3
3Kimi K2.548.3
4Qwen 3.5 397B A17B47.8
5GPT-OSS-120B34.8
6MiniMax-M2.531
7Grok 4.1 Fast (Reasoning)30.8

Interactive version: theaggregate.ai/benchmark?slug=livemathematicianbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.