CanItEdit (Lazy): leaderboard
Metric: pass@1 (%; share of the 105 hand-written Python editing problems whose edited code passes the hidden test suite, estimated from 20 samples per problem at temperature 0.2, top-p 0.95; lazy instructions, a terse user-style request). Source: arxiv.org. Saturation forecast: Estimated already saturated. 25 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4 (0613) | 51.95 |
| 2 | GPT-3.5 Turbo (0125) | 42.71 |
| 3 | CodeLlama-70B-Instruct-hf | 37.52 |
| 4 | deepseek-coder-6.7B-instruct | 31.65 |
| 5 | starcoder2-15B | 31.48 |
| 6 | deepseek-coder-6.7B-base | 27.76 |
| 7 | starcoder | 27.62 |
| 8 | Mixtral 8x7B Instruct (v0.1) | 24.9 |
| 9 | CodeLlama-34B-Instruct-hf | 24.15 |
| 10 | CodeLlama-7B-Instruct-hf | 23.49 |
| 11 | deepseek-coder-1.3B-instruct | 17.27 |
| 12 | CodeLlama-13B-Instruct-hf | 16.89 |
| 13 | starcoder2-7B | 13.76 |
| 14 | starcoder2-3B | 13.33 |
Interactive version: theaggregate.ai/benchmark?slug=canitedit-lazy · How It Works · Data refreshed daily, snapshot 2026-09-26.