NL2AGBench: leaderboard

Metric: Executable translation rate (%; share of 48 Olympiad-style geometry problems from the JGEX collection with verified AlphaGeometry formalizations whose English statement the model translates, zero-shot with the AlphaGeometry syntax definitions and handwritten clause rules in the prompt, into a specification that AlphaGeometry executes without error). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1GPT-5.475
2Grok 352.08
3Claude Sonnet 4.637.5

Interactive version: theaggregate.ai/benchmark?slug=nl2agbench · How It Works · Data refreshed daily, snapshot 2026-09-26.