EvoEval Concise — leaderboard

Metric: Pass@1 (%). Source: evo-eval.github.io. 51 models tracked.

Top models

#ModelScore
1GPT-481.1
2GPT-4 Turbo79.27
3Claude 3 Haiku71.95
4deepseek-coder-6.7B-instruct71.95
5Claude 269.51
6deepseek-coder-1.3B-instruct60.98
7Gemini 1.0 Pro59.76
8Phi-250.61
9Mixtral 8x7B Instruct43.9
10Qwen-14B43.29
11starcoder2-15B42.07
12starcoder2-7B35.37
13starcoder33.54
14starcoder2-3B33.54
15gemma-7B31.1

Interactive version: theaggregate.ai/benchmark?slug=evoeval-concise · How the rankings work · Data refreshed daily, snapshot 2026-07-22.