BigCode Models Leaderboard — leaderboard

Comprehensive leaderboard evaluating code generation models on HumanEval, MultiPL-E, DS-1000, and related benchmarks. Covers 60+ models including StarCoder, CodeLlama, Qwen-Coder, and DeepSeek-Coder variants.

Metric: HumanEval Python Pass@1 (%). Source: huggingface.co. Status: saturated. 60 models tracked.

Top models

#ModelScore
1Nxcode-CQ-7B-orpo87.2
2Qwen 2.5 Coder 32B Instruct83.2
3WizardCoder-Python-34B-V1.070.7
4Qwen 2.5 Coder 32B57.1
5deepseek-coder-33B-base52.5
6Phi-151.2
7CodeQwen1.5-7B50.8
8CodeLlama-34B-Instruct50.8
9starcoder2-15B44.1
10falcon-180B35.4
11starcoder2-7B34.1
12starcoder2-3B31.4

Interactive version: theaggregate.ai/benchmark?slug=bigcode-models-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.