MBPP+ — leaderboard

MBPP+ evaluates model capability on coding & software engineering tasks from the linked upstream source with MBPP+ pass@1 as the primary reported metric.

Metric: MBPP+ pass@1 (self-reported). Source: benchmarklist.com. Status: saturation imminent. 25 models tracked.

Top models

#ModelScore
1O1 Preview80.2
2O1 Mini78.8
3Qwen 2.5 Coder 32B Instruct77
4Gemini 1.5 Pro74.6
5Claude 3.5 Sonnet (20240620)74.3
6Claude 3 Opus73.3
7GPT-4 Turbo73.3
8DeepSeek V373
9GPT-4o72.2
10GPT-4o Mini72.2
11OpenCoder-8B-Instruct71.4
12GPT-3.5 Turbo69.7
13Claude 3 Sonnet69.3
14Llama 3 70B Instruct69
15Claude 3 Haiku68.8

Interactive version: theaggregate.ai/benchmark?slug=mbpp-plus · How the rankings work · Data refreshed daily, snapshot 2026-07-22.