CanItEdit (Lazy): leaderboard

Metric: pass@1 (%; share of the 105 hand-written Python editing problems whose edited code passes the hidden test suite, estimated from 20 samples per problem at temperature 0.2, top-p 0.95; lazy instructions, a terse user-style request). Source: arxiv.org. Saturation forecast: Estimated already saturated. 25 models tracked.

Top models

#ModelScore
1GPT-4 (0613)51.95
2GPT-3.5 Turbo (0125)42.71
3CodeLlama-70B-Instruct-hf37.52
4deepseek-coder-6.7B-instruct31.65
5starcoder2-15B31.48
6deepseek-coder-6.7B-base27.76
7starcoder27.62
8Mixtral 8x7B Instruct (v0.1)24.9
9CodeLlama-34B-Instruct-hf24.15
10CodeLlama-7B-Instruct-hf23.49
11deepseek-coder-1.3B-instruct17.27
12CodeLlama-13B-Instruct-hf16.89
13starcoder2-7B13.76
14starcoder2-3B13.33

Interactive version: theaggregate.ai/benchmark?slug=canitedit-lazy · How It Works · Data refreshed daily, snapshot 2026-09-26.