AA LiveCodeBench — leaderboard

Artificial Analysis independent evaluation of LiveCodeBench: contamination-free coding from LeetCode, AtCoder, and CodeForces problems.

Metric: Pass@1 (%). Source: artificialanalysis.ai. Status: saturated. 343 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview) (High)91.75
2DeepSeek V3.2 Speciale89.63
3GPT-5.2 (Medium)89.42
4GLM-4.7 (Reasoning)89.42
5GPT-5.2 (xHigh)88.89
6GPT-OSS-120B (High)87.83
7Claude Opus 4.5 (Thinking)87.09
8GPT-5.1 (High)86.77
9MiMo-V2-Flash (Reasoning)86.77
10DeepSeek V3.2 (Thinking)86.24
11O4 Mini (High)85.93
12Kimi K2 (Thinking)85.29
13GPT-5.1 Codex (High)84.87
14GPT-5 (High)84.55
15GPT-5 Codex (High)84.02

Interactive version: theaggregate.ai/benchmark?slug=aa-livecodebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.