CodeEditorBench Plus - Code Translate (Three-shot): leaderboard

Metric: pass@1 (%, greedy decoding, three-shot, CodeEditorBench_Plus). Source: codeeditorbench.github.io. Saturation forecast: Estimated already saturated. 19 models tracked.

Top models

#ModelScore
1GPT-4 (0613)51.7
2Gemini 1.0 Pro39.2
3GPT-3.5 Turbo (1106)36.4
4CodeLlama-13B-Instruct-hf32.7
5CodeLlama-34B-hf30.7
6CodeLlama-34B-Instruct-hf30.3
7GLM-429.9
8Phind-CodeLlama-34B-v227.5
9CodeLlama-7B-Instruct-hf27.1
10WizardCoder-15B-V1.027.1
11Magicoder-S-CL-7B24.5

Interactive version: theaggregate.ai/benchmark?slug=codeeditorbench-plus-code-translate-three-shot · How It Works · Data refreshed daily, snapshot 2026-09-26.