BigCodeBench — leaderboard

Challenging coding benchmark with 1,140 function-level tasks requiring composition of multiple function calls and complex instructions.

Metric: Pass@1 (%). Source: bigcode-bench.github.io. Status: saturation imminent. 126 models tracked.

Top models

#ModelScore
1Human Expert97
2GPT-4o (2024-05-13)51.1
3DeepSeek V350
4Llama 4 Maverick49.7
5Qwen 2.5 Coder 32B Instruct49
6GPT-4.1 Mini48.9
7GPT-4 Turbo48.2
8Qwen 2.5 Coder 14B Instruct48.2
9GPT-4o (2024-11-20)48
10Athene-V2-Chat47.2
11Llama 3.3 70B Instruct46.9
12Claude 3.5 Sonnet (20240620)46.8
13Llama 3.1 70B Instruct46.1
14GPT-4o Mini (2024-07-18)46.1
15Claude 3.5 Haiku (20241022)46.1

Interactive version: theaggregate.ai/benchmark?slug=bigcodebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.