FrontierSWE — leaderboard

Ultra-long-horizon coding agent benchmark (20h per task) testing implementation, performance engineering, and ML research at the edge of human ability. 17 tasks scored on a continuous 0-1 scale with 5 trials each.

Metric: Dominance (%). Source: www.frontierswe.com. Status: saturation imminent. 15 models tracked.

Top models

#ModelScore
1Claude Fable 589
2Grok 4.578
3Claude Opus 4.873
4GLM-5.272
5GPT-5.570
6Claude Opus 4.761
7Claude Opus 4.653
8GPT-5.451
9Gemini 3.1 Pro (Preview)37
10GLM-5.129
11DeepSeek V4 Pro27
12Kimi K2.525
13Kimi K2.625
14Qwen 3.6 Plus21

Interactive version: theaggregate.ai/benchmark?slug=frontierswe · How the rankings work · Data refreshed daily, snapshot 2026-07-22.