LongCLI-Bench - Pass@3: leaderboard

Metric: Pass@3 (%): share of the 20 long-horizon command-line programming tasks of LongCLI-Bench (CS course assignments and real research and engineering workflows: from scratch, feature addition, bug fix and refactor; isolated Docker environments; about 15,000 lines of code per task); each agent runs through its own harness (Codex, Claude Code or OpenHands) with a unified system prompt; passed (requirement and regression tests complete) in at least one of three attempts; higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 9 models tracked.

Top models

#ModelScoreOverall rank
1Claude Opus 4.625#60
2Claude Opus 4.520#79
3GPT-5.3 Codex20#68
4GPT-5.2 Codex15#89
5GPT-5.1 Codex Max15#94
6Claude Sonnet 4.510#138
7DeepSeek V3.110#260
8Qwen 3 235B A22B10#304
9GLM-4.610#246

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=longcli-bench-pass-3 · How It Works · Data refreshed daily, snapshot 2026-10-11.