AdvancedMathBench - ProverBench Qualifying Exam: leaderboard

Metric: Proof acceptance rate (%) on the 45 doctoral qualifying-exam problems (QE split): temperature 1.0, 64k maximum output tokens, highest available reasoning effort; each proof is accepted only if all 8 parallel passes of the authors' expert-trained automatic verifier accept it (pessimistic verification); higher is better. Source: arxiv.org. Saturation forecast: Around January 2028. 11 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)48.9
2GPT-5.5 (High)46.1
3DeepSeek V4 Pro (Max)40
4Claude Opus 4.8 (Max)40
5Qwen 3.5 397B A17B33.5
6GLM-5.2 (Max)28.9
7GPT-5.2 (xHigh)26.7
8Gemini 3.1 Pro (Preview) (High)17.8
9GPT-OSS-120B (High)2.2

Interactive version: theaggregate.ai/benchmark?slug=advancedmathbench-proverbench-qualifying-exam · How It Works · Data refreshed daily, snapshot 2026-09-29.