CanItEdit (Descriptive): leaderboard
Metric: pass@1 (%; share of the 105 hand-written Python editing problems whose edited code passes the hidden test suite, estimated from 20 samples per problem at temperature 0.2, top-p 0.95; descriptive instructions that spell out the change). Source: arxiv.org. Saturation forecast: Estimated already saturated. 25 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4 (0613) | 63.33 |
| 2 | GPT-3.5 Turbo (0125) | 48.14 |
| 3 | CodeLlama-70B-Instruct-hf | 45.05 |
| 4 | starcoder2-15B | 41.95 |
| 5 | deepseek-coder-6.7B-instruct | 41.03 |
| 6 | starcoder | 37.1 |
| 7 | CodeLlama-7B-Instruct-hf | 32.83 |
| 8 | deepseek-coder-6.7B-base | 32.62 |
| 9 | CodeLlama-34B-Instruct-hf | 30.63 |
| 10 | Mixtral 8x7B Instruct (v0.1) | 30.1 |
| 11 | CodeLlama-13B-Instruct-hf | 26.9 |
| 12 | deepseek-coder-1.3B-instruct | 26.22 |
| 13 | starcoder2-7B | 25.1 |
| 14 | starcoder2-3B | 15.95 |
Interactive version: theaggregate.ai/benchmark?slug=canitedit-descriptive · How It Works · Data refreshed daily, snapshot 2026-09-26.