SWE-Arena — leaderboard
Software Engineering Arena: head-to-head Elo-rated coding evaluation where 37 models solve real software engineering tasks, judged by pairwise human preference.
Metric: Elo Score. Source: swe-arena.com. 37 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O3 | 1002.01 |
| 2 | Claude 3.5 Sonnet | 1002 |
| 3 | DeepSeek V3.1 | 1002 |
| 4 | GLM 4.5 Air | 1002 |
| 5 | Grok 3 Mini | 1002 |
| 6 | Qwen 3 VL 30B A3B (Thinking) | 1002 |
| 7 | Mistral 7B Instruct | 1002 |
| 8 | GLM-4 32B | 1002 |
| 9 | Hermes-2-Pro-Llama-3-8B | 1002 |
| 10 | GPT-4o | 1000 |
| 11 | Gemini 2.5 Pro (Preview 06-05) | 1000 |
| 12 | Gemini 2.5 Pro | 998 |
| 13 | GPT-5 Mini | 998 |
| 14 | Claude 3.7 Sonnet | 998 |
| 15 | Gemma 3 27B | 998 |
Interactive version: theaggregate.ai/benchmark?slug=swe-arena · How the rankings work · Data refreshed daily, snapshot 2026-07-22.