TaoBench (Pass@128) - MathLib Formulation: leaderboard

Metric: Pass@128 (%): share of TaoBench's 150 Analysis I exercises translated into mathematically equivalent statements over standard MathLib definitions (GPT-5.1 translation with equivalence checking and expert review) for which at least one of 128 sampled whole proofs (temperature 1.0, up to 8,192 reasoning tokens) passes the Lean checker; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.

Top models

#ModelScoreOverall rank
1Goedel-Prover-V2-32B72.67
2Goedel-Prover-V2-8B70.67
3DeepSeek-Prover-V2-7B69.33
4DeepSeek-Prover-V2-7B (Non-reasoning)54.67
5Kimina-Prover-Distill-8B50

Interactive version: theaggregate.ai/benchmark?slug=taobench-pass-128-mathlib-formulation · How It Works · Data refreshed daily, snapshot 2026-10-11.