ENAMEL: leaderboard

Efficiency-aware code-generation benchmark built from HumanEval problems with expert efficient reference solutions and strong test generators, reporting eff@1 alongside pass@1.

Metric: eff@1 (self-reported). Source: benchmarklist.com. Status: saturated. 32 models tracked.

Top models

#ModelScore
1GPT-4 Turbo (Preview)47
2GPT-4 (0613)45.4
3Llama 3 70B Instruct42.1
4Mixtral 8x22B Instruct40.8
5Claude 3 Opus40.1
6Claude 3 Haiku38.6
7Claude 3 Sonnet34.5
8Llama 3 8B Instruct34.4
9Mixtral 8x7B Instruct26.6
10starcoder19.5
11Mistral 7B15.2
12vicuna-13B12.3
13Vicuna-7B6.1

Interactive version: theaggregate.ai/benchmark?slug=enamel · How It Works · Data refreshed daily, snapshot 2026-09-05.