Arena AI Code — leaderboard

Crowdsourced Arena AI pairwise human-preference leaderboard for code generation and coding-assistant models.

Metric: Arena ELO (self-reported). Source: benchmarklist.com. Status: saturation imminent. 64 models tracked.

Top models

#ModelScore
1Claude Opus 4.71570
2Claude Opus 4.61548
3GLM-5.11532
4Claude Sonnet 4.61526
5Muse Spark1509
6Claude Opus 4.5 (20251101) (Thinking 32K)1491
7GPT-5.5 (High)1490
8MiMo-V2.5-Pro1475
9Claude Opus 4.51467
10Qwen 3.6 Plus1465
11GPT-5.4 (High)1457
12DeepSeek V4 Pro1455
13Gemini 3.1 Pro (Preview)1454
14MiMo-V2.51446
15GLM-4.71440

Interactive version: theaggregate.ai/benchmark?slug=arena-ai-code · How the rankings work · Data refreshed daily, snapshot 2026-07-22.