CodeEditorBench Plus - Code Debug (Three-shot): leaderboard

Metric: pass@1 (%, greedy decoding, three-shot, CodeEditorBench_Plus). Source: codeeditorbench.github.io. Saturation forecast: Estimated already saturated. 19 models tracked.

Top models

#ModelScore
1GPT-4 (0613)34.5
2GPT-3.5 Turbo (1106)27
3Phind-CodeLlama-34B-v223.9
4GLM-423.3
5Gemini 1.0 Pro22.9
6CodeLlama-7B-Instruct-hf16.7
7CodeLlama-13B-Instruct-hf16
8Magicoder-S-CL-7B15.7
9CodeLlama-34B-Instruct-hf14.3
10CodeLlama-34B-hf13.3
11WizardCoder-15B-V1.011.4

Interactive version: theaggregate.ai/benchmark?slug=codeeditorbench-plus-code-debug-three-shot · How It Works · Data refreshed daily, snapshot 2026-09-26.