ENAMEL — leaderboard

Efficiency-aware code-generation benchmark built from HumanEval problems with expert efficient reference solutions and strong test generators, reporting eff@1 alongside pass@1.

Metric: eff@1 (self-reported). Source: benchmarklist.com. Status: saturation imminent. 31 models tracked.

Top models

#ModelScore
1GPT-4 Turbo47
2GPT-445.4
3Llama 3 70B Instruct42.1
4Claude 3 Opus40.1
5Claude 3 Haiku38.6
6Claude 3 Sonnet34.5
7Llama 3 8B Instruct34.4
8starcoder19.5
9Mistral 7B15.2
10vicuna-13B12.3
11gpt-j-6B8.3
12Vicuna-7B6.1

Interactive version: theaggregate.ai/benchmark?slug=enamel · How the rankings work · Data refreshed daily, snapshot 2026-07-22.