T2 Theorem Testing: leaderboard

Metric: Testing accuracy (%): share of T2's 2,206 problems from five Lean 4 repositories in which the generated theorem, substituted for the original, lets every dependent successor theorem compile; zero-shot from a natural-language statement (written by Claude Sonnet 4.5) with predecessor and successor context, one completion at temperature 0.6, top-p 0.95; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 18 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.538.9
2GPT-5 Mini37.9
3GPT-537.7
4Llama 3.1 70B37
5Claude 3.7 Sonnet36.8
6GPT-4o Mini36.7
7GPT-5 Nano36.6
8DeepSeek R136.3
9Claude Sonnet 436
10Llama 3.1 405B33.2
11GPT-OSS-120B32.3
12Llama 3.1 8B24.8

Interactive version: theaggregate.ai/benchmark?slug=t2-theorem-testing · How It Works · Data refreshed daily, snapshot 2026-10-07.