AtomicCommitBench: leaderboard

Metric: Adjusted Rand index against the observed developer commit history (-1 to 1, 0 for chance agreement; mean over 800 squashed multi-commit episodes from 10 Python projects, grouping the hunks of each squashed diff into replayable commits; one run per episode, at most 50 turns). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1GPT-5.40.46
2GLM-50.43
3MiniMax-M2.50.31
4Kimi K2.50.29

Interactive version: theaggregate.ai/benchmark?slug=atomiccommitbench · How It Works · Data refreshed daily, snapshot 2026-09-29.