HumanEval+ — leaderboard

HumanEval+ code-generation leaderboard from EvalPlus.

Metric: HumanEval+ pass@1 (self-reported). Source: benchmarklist.com. Status: saturation imminent. 24 models tracked.

Top models

#ModelScore
1O1 Mini89
2O1 Preview89
3GPT-4o87.2
4Qwen 2.5 Coder 32B Instruct87.2
5DeepSeek V386.6
6GPT-4 Turbo86.6
7GPT-4o Mini83.5
8Claude 3.5 Sonnet (20240620)81.7
9Grok Beta80.5
10Gemini 1.5 Pro79.3
11GPT-479.3
12Claude 3 Opus77.4
13OpenCoder-8B-Instruct77.4
14Gemini 1.5 Flash75.6
15Codestral-22B-v0.173.8

Interactive version: theaggregate.ai/benchmark?slug=humaneval-plus · How the rankings work · Data refreshed daily, snapshot 2026-07-22.