LLM2014 Code 2025-09 — leaderboard

Metric: Median Score. Source: raw.githubusercontent.com. 19 models tracked.

Top models

#ModelScore
1GPT-5 Mini (High)75.09
2Claude Sonnet 4 (Thinking)61.01
3Claude Sonnet 451.35
4Gemini 2.5 Pro47.77
5DeepSeek V3.1 (Thinking)47.5
6Grok 437.44
7GLM-4.5 (Thinking)37.21
8Gemini 2.5 Flash (Thinking)36.34
9GPT-OSS-20B (High)36.2
10GLM-4.535.42
11Qwen 3 235B A22B 2507 Instruct33.22
12DeepSeek V3.132.69
13Kimi K2 (0711)27.7
14Grok Code Fast 125.06
15GPT-OSS-120B (High)21.53

Interactive version: theaggregate.ai/benchmark?slug=llm2014-code-2025-09 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.