KGHaluBench - Depth Hallucination Rate: leaderboard
Metric: Depth-of-knowledge hallucination rate (%): incorrect facts found by the fact-level check (run on responses that pass the entity filter) divided by the maximum attainable score, over 10 assessments of 150 questions each (compound questions about Wikidata entities, generated dynamically from the knowledge graph), default sampling parameters; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 25 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5 | 23.87 | #91 |
| 2 | Grok 4 | 24.4 | #169 |
| 3 | Gemini 2.5 Pro | 26.97 | #145 |
| 4 | Llama 3.1 405B Instruct | 27.46 | #447 |
| 5 | GPT-5 Mini | 28.05 | #176 |
| 6 | GPT-4.1 | 28.65 | #240 |
| 7 | Claude Opus 4.1 (20250805) | 28.79 | #126 |
| 8 | Grok 3 | 29.05 | #296 |
| 9 | GLM-4.5 | 30.01 | #265 |
| 10 | Kimi K2 | 30.47 | #236 |
| 11 | Claude Sonnet 4 (20250514) | 32.55 | #211 |
| 12 | DeepSeek R1 0528 | 33.5 | #217 |
| 13 | Gemini 2.5 Flash | 34.41 | #237 |
| 14 | DeepSeek V3 (0324) | 35.9 | #332 |
| 15 | GPT-4.1 Mini | 37.71 | #346 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=kghalubench-depth-hallucination-rate · How It Works · Data refreshed daily, snapshot 2026-10-11.