SWE-Marathon — leaderboard

Long-horizon software engineering benchmark where coding agents work on realistic repository tasks under marathon-scale time budgets, reporting pass@1 for end-to-end completed tasks.

Metric: Pass@1 (%). Source: www.swe-marathon.org. Status: saturation imminent. 11 models tracked.

Top models

#ModelScore
1Claude Opus 4.826
2Claude Fable 524
3Claude Opus 4.716
4GLM-5.213
5GPT-5.512
6Gemini 3.5 Flash7
7Gemini 3.1 Pro (Preview)4
8DeepSeek V4 Pro4
9GLM-5.11
10Kimi K2.60
11MiniMax-M2.70

Interactive version: theaggregate.ai/benchmark?slug=swe-marathon · How the rankings work · Data refreshed daily, snapshot 2026-07-22.