Delulu - Hallucination Detection: leaderboard

Metric: Both-correct rate (%) on all 1,951 Delulu samples as an LLM judge: the judge sees each sample twice, once with the gold completion and once with the hallucinated one, and must accept the first and reject the second (both-correct rate, %), with chain-of-thought and a binary verdict; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.592.1
2GPT-5.2 Codex88.2
3Claude Sonnet 4.685.8
4GPT-5.482.8
5Claude Haiku 4.577.9
6GPT-5.4 Mini74.8
7GPT-4.1 Mini67.7
8GPT-4o Mini52.3

Interactive version: theaggregate.ai/benchmark?slug=delulu-hallucination-detection · How It Works · Data refreshed daily, snapshot 2026-10-07.