Delulu - Parameter Hallucination Detection: leaderboard

Metric: Both-correct rate (%) on the 435 Delulu samples whose hallucination is parameters that do not exist, as an LLM judge: the judge sees each sample twice, once with the gold completion and once with the hallucinated one, and must accept the first and reject the second (both-correct rate, %), with chain-of-thought and a binary verdict; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.591.1
2Claude Sonnet 4.688.5
3GPT-5.2 Codex87.2
4GPT-5.479.8
5Claude Haiku 4.574.3
6GPT-5.4 Mini66.3
7GPT-4.1 Mini63.5
8GPT-4o Mini46.1

Interactive version: theaggregate.ai/benchmark?slug=delulu-parameter-hallucination-detection · How It Works · Data refreshed daily, snapshot 2026-10-07.