LongCLI-Bench - Fail-to-Pass Rate: leaderboard

Metric: Fail-to-pass rate (%), mean of three attempts: share of the 20 long-horizon command-line programming tasks of LongCLI-Bench (CS course assignments and real research and engineering workflows: from scratch, feature addition, bug fix and refactor; isolated Docker environments; about 15,000 lines of code per task); each agent runs through its own harness (Codex, Claude Code or OpenHands) with a unified system prompt; whose requirement (fail-to-pass) tests all pass, regressions aside; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScoreOverall rank
1Claude Opus 4.620#60
2Claude Opus 4.520#79
3GPT-5.3 Codex18.3#68
4GPT-5.1 Codex Max16.7#94
5Claude Sonnet 4.513.3#138
6Qwen 3 235B A22B13.3#304
7GPT-5.2 Codex13.3#89
8DeepSeek V3.111.7#260
9GLM-4.611.7#246

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=longcli-bench-fail-to-pass-rate · How It Works · Data refreshed daily, snapshot 2026-10-11.