NL2AGBench: leaderboard
Metric: Executable translation rate (%; share of 48 Olympiad-style geometry problems from the JGEX collection with verified AlphaGeometry formalizations whose English statement the model translates, zero-shot with the AlphaGeometry syntax definitions and handwritten clause rules in the prompt, into a specification that AlphaGeometry executes without error). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 75 |
| 2 | Grok 3 | 52.08 |
| 3 | Claude Sonnet 4.6 | 37.5 |
Interactive version: theaggregate.ai/benchmark?slug=nl2agbench · How It Works · Data refreshed daily, snapshot 2026-09-26.