KGHaluBench - Breadth Hallucination Rate: leaderboard

Metric: Breadth-of-knowledge hallucination rate (%): share of non-abstained responses that the entity-level filter classifies as describing the wrong entity, over 10 assessments of 150 questions each (compound questions about Wikidata entities, generated dynamically from the knowledge graph), default sampling parameters; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 25 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Mini7.69#176
2GPT-58.2#91
3Grok 410.27#169
4Claude Opus 4.1 (20250805)10.32#126
5GPT-4.110.84#240
6Llama 3.1 405B Instruct10.88#447
7Claude Sonnet 4 (20250514)13.01#211
8Gemini 2.5 Pro14.18#145
9Grok 315.2#296
10GLM-4.516.24#265
11Qwen 3 235B A22B 2507 Instruct21.49#291
12Gemini 2.5 Flash21.55#237
13GPT-4.1 Mini21.62#346
14GPT-OSS-120B22.74#330
15Kimi K223.66#236

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=kghalubench-breadth-hallucination-rate · How It Works · Data refreshed daily, snapshot 2026-10-11.