Natural2Code: leaderboard

Google's internal held-out Python code-generation set in HumanEval format, author-written to avoid web leakage, reported 0-shot in the Gemini 1.0 report (2023): Gemini Ultra 74.9%, GPT-4 73.9%.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturated. 24 models tracked.

Top models

#ModelScore
1GPT-452.8
2GPT-4 Turbo (Preview)51.5
3Claude 3 Opus48.3
4Gemini 1.5 Pro42.3
5GPT-3.5 Turbo40.7
6Claude 3 Sonnet38.9
7Claude 3 Haiku36.2
8Claude 2.134.4
9DeepSeek V3 Chat33.2
10starcoder13.2

Interactive version: theaggregate.ai/benchmark?slug=natural2code · How It Works · Data refreshed daily, snapshot 2026-09-05.