ProofBench: leaderboard

200 Lean 4 proof problems (100 public, 100 private) from graduate qualifying exams and textbooks in analysis, algebra, number theory and logic; Vals AI, 2026; pass only if it compiles without sorry.

Metric: Accuracy (%). Source: www.vals.ai. Status: saturation imminent. 31 models tracked.

Top models

#ModelScore
1Claude Fable 5.1 (Max)100
2Claude Opus 5 (Max)99
3GPT-6 (Max)99
4Claude Fable 5 (Max)95
5Kimi K387
6GPT-5.6 Sol (Max)83
7Claude Sonnet 5 (Max)77
8GPT-5.6 Terra (xHigh)74
9GPT-5.6 Terra (Max)74
10GPT-5.6 Luna (Max)60
11Qwen 3.8 Max58
12Gemini 3.7 Flash (High)58
13DeepSeek V4 Flash (0731)56
14Grok 4.6 (High)51
15DeepSeek V4 Pro (0813) (Max)50

Interactive version: theaggregate.ai/benchmark?slug=proofbench · How It Works · Data refreshed daily, snapshot 2026-09-05.