Leipzig Benchmark: leaderboard

Metric: Questions solved (out of 100) in a single attempt per question: 100 research-level mathematics questions with unique, unguessable answers from 49 mathematicians (Benchmarks in Leipzig, MPI MiS, 2026), API runs with web search and code tools disabled, errors or timeouts rerun up to three times, answers checked by GPT-5.5 and Gemini 3.1 Pro judges; questions were accepted only if at most three of the five project-active models solved them; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 10 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)44
2GPT-5.428
3Gemini 3.1 Pro (Preview) (Medium)15
4Claude Opus 4.614
5Gemini 3 Pro14
6Claude Opus 4.713
7DeepSeek V4 Pro (High)10
8DeepSeek V3.28
9Grok 4.3 (High)6
10Grok 4.205

Interactive version: theaggregate.ai/benchmark?slug=leipzig-benchmark · How It Works · Data refreshed daily, snapshot 2026-09-29.