Scale AI - SWE-Bench Pro: leaderboard

Extended SWE-bench testing long-horizon tasks on both public open-source and private commercial repositories. Part of Scale's SEAL evaluation platform.

Metric: Score. Source: scale.com. Status: years away from saturation. 25 models tracked.

Top models

#ModelScore
1Muse Spark 1.161.5
2GPT-5.4 (xHigh)59.1
3Muse Spark55
4Claude Opus 4.6 (Thinking)51.9
5Gemini 3.1 Pro (Preview) (Thinking)46.1
6Claude Opus 4.5 (20251101)45.89
7Qwen3 Coder Next44.3
8Claude Sonnet 4.543.6
9Gemini 3 Pro (Preview)43.3
10Claude Sonnet 442.7
11GPT-5 (High)41.78
12GPT-5.2 Codex41.04
13Claude Haiku 4.539.45
14Qwen 3 Coder 480B A35B Instruct38.7
15MiniMax-M2.136.81

Interactive version: theaggregate.ai/benchmark?slug=scale-ai-swe-bench-pro · How It Works · Data refreshed daily, snapshot 2026-09-05.