MBPP+: leaderboard

378 hand-verified Python problems drawn from MBPP-sanitized, each with 35x more test cases than the original MBPP, built by the EvalPlus project (NeurIPS 2023); pass@1 with greedy decoding.

Metric: MBPP+ pass@1 (self-reported). Source: benchmarklist.com. Status: saturated. 25 models tracked.

Top models

#ModelScore
1O1 Mini (2024-09-12)78.8
2Qwen 2.5 Coder 32B Instruct77
3Gemini 1.5 Pro (002)74.6
4Claude 3.5 Sonnet (20240620)74.3
5Claude 3 Opus73.3
6GPT-4 Turbo (Preview)73.3
7GPT-4o Mini (2024-07-18)72.2
8GPT-4o (2024-08-06)72.2
9OpenCoder-8B-Instruct71.4
10GPT-3.5 Turbo (1106)69.7
11Claude 3 Sonnet69.3
12Llama 3 70B Instruct69
13Claude 3 Haiku68.8
14Gemini 1.5 Flash (002)67.5

Interactive version: theaggregate.ai/benchmark?slug=mbpp-plus · How It Works · Data refreshed daily, snapshot 2026-09-05.