TLA+-Bench: leaderboard

Metric: TLC-correct rate (%; default regime, the model must recover the interface names) on 100 model-checked gold TLA+ specifications (system category, trivial single-state fixtures removed), written from a GPT-5 declarative description, one sample each; correct when TLC model-checks the output against the reference configuration. Source: arxiv.org. Saturation forecast: Around 2029. 6 models tracked.

Top models

#ModelScore
1Claude Opus 4.516
2Gemini 2.5 Pro10
3GPT-54

Interactive version: theaggregate.ai/benchmark?slug=tla-plus-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.