Delulu - Exact Match: leaderboard

Metric: Exact match (%) of the completion with the gold completion, byte for byte, on all 1,951 Delulu fill-in-the-middle samples (seven languages, four injected hallucination types), greedy decoding, at most 256 tokens; base FIM models are truncated to the gold completion's line count; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1Qwen 2.5 Coder 32B Instruct54.1
2CodeLlama-13B-hf47.3
3CodeLlama-7B-hf46.8
4Qwen 2.5 Coder 14B Instruct46.2
5Qwen 2.5 Coder 3B Instruct44
6Qwen 2.5 Coder 7B Instruct40.2
7Qwen2.5-Coder-1.5B-Instruct37.8
8starcoder2-15B21.5
9starcoder2-7B9.7

Interactive version: theaggregate.ai/benchmark?slug=delulu-exact-match · How It Works · Data refreshed daily, snapshot 2026-10-07.