CanItEdit (Descriptive): leaderboard

Metric: pass@1 (%; share of the 105 hand-written Python editing problems whose edited code passes the hidden test suite, estimated from 20 samples per problem at temperature 0.2, top-p 0.95; descriptive instructions that spell out the change). Source: arxiv.org. Saturation forecast: Estimated already saturated. 25 models tracked.

Top models

#ModelScore
1GPT-4 (0613)63.33
2GPT-3.5 Turbo (0125)48.14
3CodeLlama-70B-Instruct-hf45.05
4starcoder2-15B41.95
5deepseek-coder-6.7B-instruct41.03
6starcoder37.1
7CodeLlama-7B-Instruct-hf32.83
8deepseek-coder-6.7B-base32.62
9CodeLlama-34B-Instruct-hf30.63
10Mixtral 8x7B Instruct (v0.1)30.1
11CodeLlama-13B-Instruct-hf26.9
12deepseek-coder-1.3B-instruct26.22
13starcoder2-7B25.1
14starcoder2-3B15.95

Interactive version: theaggregate.ai/benchmark?slug=canitedit-descriptive · How It Works · Data refreshed daily, snapshot 2026-09-26.