Arena AI Code — leaderboard
Crowdsourced Arena AI pairwise human-preference leaderboard for code generation and coding-assistant models.
Metric: Arena ELO (self-reported). Source: benchmarklist.com. Status: saturation imminent. 64 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.7 | 1570 |
| 2 | Claude Opus 4.6 | 1548 |
| 3 | GLM-5.1 | 1532 |
| 4 | Claude Sonnet 4.6 | 1526 |
| 5 | Muse Spark | 1509 |
| 6 | Claude Opus 4.5 (20251101) (Thinking 32K) | 1491 |
| 7 | GPT-5.5 (High) | 1490 |
| 8 | MiMo-V2.5-Pro | 1475 |
| 9 | Claude Opus 4.5 | 1467 |
| 10 | Qwen 3.6 Plus | 1465 |
| 11 | GPT-5.4 (High) | 1457 |
| 12 | DeepSeek V4 Pro | 1455 |
| 13 | Gemini 3.1 Pro (Preview) | 1454 |
| 14 | MiMo-V2.5 | 1446 |
| 15 | GLM-4.7 | 1440 |
Interactive version: theaggregate.ai/benchmark?slug=arena-ai-code · How the rankings work · Data refreshed daily, snapshot 2026-07-22.