EvoEval Difficult — leaderboard

Metric: Pass@1 (%). Source: evo-eval.github.io. 51 models tracked.

Top models

#ModelScore
1GPT-452
2GPT-4 Turbo50
3Claude 3 Haiku40
4deepseek-coder-6.7B-instruct40
5Gemini 1.0 Pro37
6Claude 229
7Qwen-14B23
8Mixtral 8x7B Instruct21
9deepseek-coder-1.3B-instruct20
10Phi-218
11starcoder2-15B16
12gemma-7B12
13starcoder12
14starcoder2-7B12
15Qwen-7B9

Interactive version: theaggregate.ai/benchmark?slug=evoeval-difficult · How the rankings work · Data refreshed daily, snapshot 2026-07-22.