KGHaluBench - Depth Hallucination Rate: leaderboard

Metric: Depth-of-knowledge hallucination rate (%): incorrect facts found by the fact-level check (run on responses that pass the entity filter) divided by the maximum attainable score, over 10 assessments of 150 questions each (compound questions about Wikidata entities, generated dynamically from the knowledge graph), default sampling parameters; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 25 models tracked.

Top models

#ModelScoreOverall rank
1GPT-523.87#91
2Grok 424.4#169
3Gemini 2.5 Pro26.97#145
4Llama 3.1 405B Instruct27.46#447
5GPT-5 Mini28.05#176
6GPT-4.128.65#240
7Claude Opus 4.1 (20250805)28.79#126
8Grok 329.05#296
9GLM-4.530.01#265
10Kimi K230.47#236
11Claude Sonnet 4 (20250514)32.55#211
12DeepSeek R1 052833.5#217
13Gemini 2.5 Flash34.41#237
14DeepSeek V3 (0324)35.9#332
15GPT-4.1 Mini37.71#346

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=kghalubench-depth-hallucination-rate · How It Works · Data refreshed daily, snapshot 2026-10-11.