KGHaluBench: leaderboard
Metric: Weighted accuracy in percent points: points earned through correct facts and abstentions, scaled by the assessment's estimated question difficulty relative to the average difficulty (so it can in principle exceed 100), over 10 assessments of 150 questions each (compound questions about Wikidata entities, generated dynamically from the knowledge graph), default sampling parameters; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 25 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5 | 65.6 | #91 |
| 2 | Grok 4 | 63.66 | #169 |
| 3 | GPT-4.1 | 63.6 | #240 |
| 4 | Gemini 2.5 Pro | 61.27 | #145 |
| 5 | Grok 3 | 57.24 | #296 |
| 6 | GLM-4.5 | 54.35 | #265 |
| 7 | Kimi K2 | 52.88 | #236 |
| 8 | Claude Opus 4.1 (20250805) | 52.71 | #126 |
| 9 | Llama 3.1 405B Instruct | 52.06 | #447 |
| 10 | Gemini 2.5 Flash | 51.41 | #237 |
| 11 | GPT-5 Mini | 49.58 | #176 |
| 12 | DeepSeek R1 0528 | 48.45 | #217 |
| 13 | GPT-4.1 Mini | 48.13 | #346 |
| 14 | DeepSeek V3 (0324) | 47 | #332 |
| 15 | Claude Sonnet 4 (20250514) | 46.18 | #211 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=kghalubench · How It Works · Data refreshed daily, snapshot 2026-10-11.