TLA+-Bench (Configuration-Aware): leaderboard
Metric: TLC-correct rate (%; configuration-aware regime, the prompt also supplies the constant and property names the configuration uses; a separate set of generations) on 100 model-checked gold TLA+ specifications (system category, trivial single-state fixtures removed), written from a GPT-5 declarative description, one sample each; correct when TLC model-checks the output against the reference configuration. Source: arxiv.org. Saturation forecast: Around May 2028. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.5 | 26 |
| 2 | Gemini 2.5 Pro | 19 |
| 3 | GPT-5 | 11 |
Interactive version: theaggregate.ai/benchmark?slug=tla-plus-bench-configuration-aware · How It Works · Data refreshed daily, snapshot 2026-09-29.