MathAtlas - Statements: leaderboard

Metric: Correctness (%; share of all theorem, example and exercise statements whose Lean 4 formalization compiles against Mathlib (Lean v4.24.0) and is judged faithful to the informal text by CriticLean-32B; zero-shot prompt for the general models, each fine-tuned formalizer's own prompt). Source: arxiv.org. Saturation forecast: Around 2028. 7 models tracked.

Top models

#ModelScore
1GPT-OSS-120B6.8
2GPT-OSS-20B4.2

Interactive version: theaggregate.ai/benchmark?slug=mathatlas-statements · How It Works · Data refreshed daily, snapshot 2026-09-26.