Delulu - CodeBLEU: leaderboard

Metric: CodeBLEU (0-100): n-gram, syntax-tree and dataflow agreement with the gold completion, on all 1,951 Delulu fill-in-the-middle samples (seven languages, four injected hallucination types), greedy decoding, at most 256 tokens; base FIM models are truncated to the gold completion's line count; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1Qwen 2.5 Coder 32B Instruct53.1
2Qwen 2.5 Coder 14B Instruct51.5
3CodeLlama-13B-hf47.8
4Qwen 2.5 Coder 7B Instruct47.2
5Qwen 2.5 Coder 3B Instruct47.2
6starcoder2-15B46.6
7CodeLlama-7B-hf46.3
8starcoder2-7B44.5
9Qwen2.5-Coder-1.5B-Instruct42.3

Interactive version: theaggregate.ai/benchmark?slug=delulu-codebleu · How It Works · Data refreshed daily, snapshot 2026-10-07.