Delulu - Method Hallucination Detection: leaderboard

Metric: Both-correct rate (%) on the 461 Delulu samples whose hallucination is calls to methods that do not exist, as an LLM judge: the judge sees each sample twice, once with the gold completion and once with the hallucinated one, and must accept the first and reject the second (both-correct rate, %), with chain-of-thought and a binary verdict; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.594.6
2GPT-5.2 Codex92.7
3GPT-5.484.2
4Claude Sonnet 4.678.4
5GPT-5.4 Mini78.2
6Claude Haiku 4.577.1
7GPT-4.1 Mini73
8GPT-4o Mini54.6

Interactive version: theaggregate.ai/benchmark?slug=delulu-method-hallucination-detection · How It Works · Data refreshed daily, snapshot 2026-10-07.