TLA+-Bench: leaderboard
Metric: TLC-correct rate (%; default regime, the model must recover the interface names) on 100 model-checked gold TLA+ specifications (system category, trivial single-state fixtures removed), written from a GPT-5 declarative description, one sample each; correct when TLC model-checks the output against the reference configuration. Source: arxiv.org. Saturation forecast: Around 2029. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.5 | 16 |
| 2 | Gemini 2.5 Pro | 10 |
| 3 | GPT-5 | 4 |
Interactive version: theaggregate.ai/benchmark?slug=tla-plus-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.