TaoBench (Pass@128) - MathLib Formulation: leaderboard
Metric: Pass@128 (%): share of TaoBench's 150 Analysis I exercises translated into mathematically equivalent statements over standard MathLib definitions (GPT-5.1 translation with equivalence checking and expert review) for which at least one of 128 sampled whole proofs (temperature 1.0, up to 8,192 reasoning tokens) passes the Lean checker; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Goedel-Prover-V2-32B | 72.67 | |
| 2 | Goedel-Prover-V2-8B | 70.67 | |
| 3 | DeepSeek-Prover-V2-7B | 69.33 | |
| 4 | DeepSeek-Prover-V2-7B (Non-reasoning) | 54.67 | |
| 5 | Kimina-Prover-Distill-8B | 50 |
Interactive version: theaggregate.ai/benchmark?slug=taobench-pass-128-mathlib-formulation · How It Works · Data refreshed daily, snapshot 2026-10-11.