FrontierSWE: leaderboard

Ultra-long-horizon coding agent benchmark (20h per task) testing implementation, performance engineering, and ML research at the edge of human ability. 17 tasks scored on a continuous 0-1 scale with 5 trials each.

Metric: Dominance (%). Source: www.frontierswe.com. Status: saturation imminent. 17 models tracked.

Top models

#ModelScore
1Claude Fable 588
2GLM-5.378
3Grok 4.678
4Grok 4.572
5Claude Opus 4.867
6GLM-5.267
7GPT-5.565
8Claude Opus 4.756
9Claude Opus 4.649
10GPT-5.446
11Gemini 3.1 Pro (Preview)34
12Composer 2.534
13GLM-5.126
14DeepSeek V4 Pro25
15Kimi K2.523

Interactive version: theaggregate.ai/benchmark?slug=frontierswe · How It Works · Data refreshed daily, snapshot 2026-09-05.