Delulu - Import Hallucination Detection: leaderboard

Metric: Both-correct rate (%) on the 478 Delulu samples whose hallucination is imports of packages that do not exist, as an LLM judge: the judge sees each sample twice, once with the gold completion and once with the hallucinated one, and must accept the first and reject the second (both-correct rate, %), with chain-of-thought and a binary verdict; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.583.4
2Claude Sonnet 4.680.3
3GPT-5.2 Codex75.4
4GPT-5.4 Mini69.2
5GPT-5.467.9
6Claude Haiku 4.565.6
7GPT-4.1 Mini56.9
8GPT-4o Mini48.7

Interactive version: theaggregate.ai/benchmark?slug=delulu-import-hallucination-detection · How It Works · Data refreshed daily, snapshot 2026-10-07.