EvalPlus — leaderboard

EvalPlus leaderboard aggregating HumanEval+ and MBPP+ code-generation pass@1 scores.

Metric: EvalPlus Avg. (self-reported). Source: benchmarklist.com. Status: saturation imminent. 25 models tracked.

Top models

#ModelScore
1O1 Preview84.6
2O1 Mini83.9
3Qwen 2.5 Coder 32B Instruct82.1
4DeepSeek V379.8
5GPT-4o79.7
6Claude 3.5 Sonnet (20240620)78
7GPT-4o Mini77.85
8GPT-4 Turbo77.5
9Gemini 1.5 Pro76.95
10Claude 3 Opus75.35
11OpenCoder-8B-Instruct74.4
12Grok Beta73.05
13Gemini 1.5 Flash71.55
14Llama 3 70B Instruct70.5
15GPT-3.5 Turbo70.2

Interactive version: theaggregate.ai/benchmark?slug=evalplus · How the rankings work · Data refreshed daily, snapshot 2026-07-22.