BigCodeBench-Hard — leaderboard

BigCodeBench-Hard evaluates code generation on the harder BigCodeBench subset, reporting pass@1 in complete and instruct settings.

Metric: Instruct pass@1 (self-reported). Source: benchmarklist.com. Status: saturation imminent. 20 models tracked.

Top models

#ModelScore
1O3 Mini33.1
2Claude 3.7 Sonnet32.4
3O132.4
4GPT-4.1 Mini31.8
5GPT-4.131.8
6Grok 3 Mini Beta31.1
7QwQ-32B29.7
8GPT-4 Turbo29.1
9Llama 3.3 70B Instruct28.4
10GPT-4.1 Nano28.4
11DeepSeek V328.4
12GPT-4o (2024-11-20)27.7
13Qwen 2.5 Coder 32B Instruct27.7
14O1 Mini (2024-09-12)27.7

Interactive version: theaggregate.ai/benchmark?slug=bigcodebench-hard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.