LongCLI-Bench - Fail-to-Pass Step Score: leaderboard

Metric: Fail-to-pass step score (%), mean of three attempts: mean share of requirement test steps completed over the 20 long-horizon command-line programming tasks of LongCLI-Bench (CS course assignments and real research and engineering workflows: from scratch, feature addition, bug fix and refactor; isolated Docker environments; about 15,000 lines of code per task); each agent runs through its own harness (Codex, Claude Code or OpenHands) with a unified system prompt; higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 9 models tracked.

Top models

#ModelScoreOverall rank
1Claude Opus 4.650.7#60
2Claude Opus 4.547.1#79
3GPT-5.3 Codex44.1#68
4Claude Sonnet 4.542#138
5GPT-5.1 Codex Max41.8#94
6GPT-5.2 Codex39.1#89
7Qwen 3 235B A22B28.9#304
8GLM-4.626.8#246
9DeepSeek V3.125.3#260

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=longcli-bench-fail-to-pass-step-score · How It Works · Data refreshed daily, snapshot 2026-10-11.