Delulu - Edit Similarity: leaderboard

Metric: Edit similarity (0-100): character-level normalized Levenshtein similarity to the gold completion, on all 1,951 Delulu fill-in-the-middle samples (seven languages, four injected hallucination types), greedy decoding, at most 256 tokens; base FIM models are truncated to the gold completion's line count; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1Qwen 2.5 Coder 32B Instruct76.5
2Qwen 2.5 Coder 14B Instruct71.6
3CodeLlama-13B-hf70.1
4CodeLlama-7B-hf69.5
5Qwen 2.5 Coder 3B Instruct69.3
6Qwen2.5-Coder-1.5B-Instruct63.2
7Qwen 2.5 Coder 7B Instruct62.5
8starcoder2-15B60
9starcoder2-7B59.6

Interactive version: theaggregate.ai/benchmark?slug=delulu-edit-similarity · How It Works · Data refreshed daily, snapshot 2026-10-07.