LiveMathematicianBench: leaderboard

Continuously updated benchmark testing LLMs on understanding mathematical theorems from newly published arXiv preprints. Multiple-choice questions constructed from theorem statements and proof sketches.

Metric: Accuracy (%). Source: livemathematicianbench.github.io. Status: years away from saturation. 19 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash49.5
2GPT-5.446.6
3Claude Opus 545.1
4Gemini 3.1 Pro (Preview)42.9
5DeepSeek V4 Pro36.9
6Mistral Large 333
7GPT-532.2
8Gemini 2.5 Pro32.1
9DeepSeek V3.230.8
10DeepSeek R130.6
11GPT-4.130.5
12Llama 4 Maverick29.7
13Llama 3.3 70B Instruct27
14O3 (2025-04-16)27
15GPT-5.226.3

Interactive version: theaggregate.ai/benchmark?slug=livemathematicianbench · How It Works · Data refreshed daily, snapshot 2026-09-05.