Leipzig Benchmark: leaderboard
Metric: Questions solved (out of 100) in a single attempt per question: 100 research-level mathematics questions with unique, unguessable answers from 49 mathematicians (Benchmarks in Leipzig, MPI MiS, 2026), API runs with web search and code tools disabled, errors or timeouts rerun up to three times, answers checked by GPT-5.5 and Gemini 3.1 Pro judges; questions were accepted only if at most three of the five project-active models solved them; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 (xHigh) | 44 |
| 2 | GPT-5.4 | 28 |
| 3 | Gemini 3.1 Pro (Preview) (Medium) | 15 |
| 4 | Claude Opus 4.6 | 14 |
| 5 | Gemini 3 Pro | 14 |
| 6 | Claude Opus 4.7 | 13 |
| 7 | DeepSeek V4 Pro (High) | 10 |
| 8 | DeepSeek V3.2 | 8 |
| 9 | Grok 4.3 (High) | 6 |
| 10 | Grok 4.20 | 5 |
Interactive version: theaggregate.ai/benchmark?slug=leipzig-benchmark · How It Works · Data refreshed daily, snapshot 2026-09-29.