Vellum - HumanEval: leaderboard

Vellum's independently-run HumanEval coding evaluation. Function completion tasks measuring code generation ability.

Metric: Pass@1 (%). Source: www.vellum.ai. Status: years away from saturation. 47 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol96.2
2Claude Mythos 595.5
3Claude Fable 595
4GPT-5.6 Luna93
5Claude Opus 4.888.6
6Claude Opus 4.787.6
7Claude Sonnet 585.2
8Claude Sonnet 4.582
9Claude Opus 4.580.9
10Claude Opus 4.680.8
11Gemini 3.1 Pro (Preview)80.6
12DeepSeek V4 Pro80.6
13MiniMax-M380.5
14Kimi K2.680.2
15GPT-5.280

Interactive version: theaggregate.ai/benchmark?slug=vellum-humaneval · How It Works · Data refreshed daily, snapshot 2026-09-05.