TaoBench (Pass@128): leaderboard
Metric: Pass@128 (%): share of TaoBench's 150 exercises from Terence Tao's Lean 4 formalization of Analysis I, stated in Tao's own definitional framework with the minimal compilable local context (definitions, notation, lemmas) retrieved for each for which at least one of 128 sampled whole proofs (temperature 1.0, up to 8,192 reasoning tokens) passes the Lean checker; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Goedel-Prover-V2-32B | 49.33 | |
| 2 | DeepSeek-Prover-V2-7B | 41.33 | |
| 3 | DeepSeek-Prover-V2-7B (Non-reasoning) | 38 | |
| 4 | Goedel-Prover-V2-8B | 37.33 | |
| 5 | Kimina-Prover-Distill-8B | 22 |
Interactive version: theaggregate.ai/benchmark?slug=taobench-pass-128 · How It Works · Data refreshed daily, snapshot 2026-10-11.