TLA+-Bench (Configuration-Aware): leaderboard

Metric: TLC-correct rate (%; configuration-aware regime, the prompt also supplies the constant and property names the configuration uses; a separate set of generations) on 100 model-checked gold TLA+ specifications (system category, trivial single-state fixtures removed), written from a GPT-5 declarative description, one sample each; correct when TLC model-checks the output against the reference configuration. Source: arxiv.org. Saturation forecast: Around May 2028. 6 models tracked.

Top models

#ModelScore
1Claude Opus 4.526
2Gemini 2.5 Pro19
3GPT-511

Interactive version: theaggregate.ai/benchmark?slug=tla-plus-bench-configuration-aware · How It Works · Data refreshed daily, snapshot 2026-09-29.