LiveCodeBench — leaderboard

Contamination-free code benchmark using fresh problems from LeetCode, AtCoder, and CodeForces. Tests code generation, self-repair, execution, and test prediction.

Metric: Pass@1 avg (%). Source: livecodebench.github.io. Status: saturated. 28 models tracked.

Top models

#ModelScore
1O4 Mini (High)87.3
2O3 (High)84.7
3O4 Mini (Medium)84.5
4DeepSeek R1 052884.4
5Gemini 2.5 Pro (Preview 06-05)84.3
6Qwen 3 235B A22B80.4
7Grok 3 Mini (High)78.1
8O3 Mini (2025-01-31) (High)77.7
9O4 Mini (Low)77.4
10Gemini 2.5 Flash (Preview 05-20)76.2
11O3 Mini (2025-01-31)75.4
12Gemini 2.5 Flash (Preview 04-17)75.1
13O3 Mini (Low)70.6
14Claude Opus 4 (Thinking)70.4
15Claude Sonnet 4 (Thinking)68.5

Interactive version: theaggregate.ai/benchmark?slug=livecodebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.