SWE Atlas (All Workflows): leaderboard

Metric: Pass@1 (%): mean per-trial pass rate over all 284 SWE Atlas tasks (124 Codebase Q&A, 90 Test Writing and 70 Refactoring tasks in real repositories, scored by programmatic checks plus must-have rubrics), three trials per task run through the Harbor framework, each model in its provider's own agent harness (Codex CLI, Claude Code, Gemini CLI) or in mini-SWE-agent; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 14 models tracked.

Top models

#ModelScore
1Claude Opus 4.7 (xHigh)38.94
2GPT-5.4 (xHigh)38
3Claude Opus 4.6 (High)31.83
4Claude Sonnet 4.6 (High)31.63
5GLM-524.03
6Gemini 3.1 Pro (Preview) (High)23.73
7Kimi K2.519.05
8Gemini 3 Flash (High)15.65
9MiniMax-M2.515.2

Interactive version: theaggregate.ai/benchmark?slug=swe-atlas-all-workflows · How It Works · Data refreshed daily, snapshot 2026-10-07.