AdvancedMathBench - VerifierBench: leaderboard

Metric: Meta-verification balanced F1 (%): a validity verdict counts only when its rationale and error localization agree with the expert annotation, 888 model-generated proof trajectories with expert labels; balanced F1 is the harmonic mean of the true positive and true negative rates; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 11 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro (Max)65.1
2GPT-5.5 (xHigh)64.9
3GPT-5.5 (High)63.6
4GLM-5.2 (Max)63.3
5Qwen 3.5 397B A17B58.5
6GPT-5.2 (xHigh)57.7
7Gemini 3.1 Pro (Preview) (High)55.2
8Claude Opus 4.8 (Max)51
9GPT-OSS-120B (High)47.9

Interactive version: theaggregate.ai/benchmark?slug=advancedmathbench-verifierbench · How It Works · Data refreshed daily, snapshot 2026-09-29.