AtomicCommitBench: leaderboard
Metric: Adjusted Rand index against the observed developer commit history (-1 to 1, 0 for chance agreement; mean over 800 squashed multi-commit episodes from 10 Python projects, grouping the hunks of each squashed diff into replayable commits; one run per episode, at most 50 turns). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 0.46 |
| 2 | GLM-5 | 0.43 |
| 3 | MiniMax-M2.5 | 0.31 |
| 4 | Kimi K2.5 | 0.29 |
Interactive version: theaggregate.ai/benchmark?slug=atomiccommitbench · How It Works · Data refreshed daily, snapshot 2026-09-29.