KGHaluBench: leaderboard

Metric: Weighted accuracy in percent points: points earned through correct facts and abstentions, scaled by the assessment's estimated question difficulty relative to the average difficulty (so it can in principle exceed 100), over 10 assessments of 150 questions each (compound questions about Wikidata entities, generated dynamically from the knowledge graph), default sampling parameters; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 25 models tracked.

Top models

#ModelScoreOverall rank
1GPT-565.6#91
2Grok 463.66#169
3GPT-4.163.6#240
4Gemini 2.5 Pro61.27#145
5Grok 357.24#296
6GLM-4.554.35#265
7Kimi K252.88#236
8Claude Opus 4.1 (20250805)52.71#126
9Llama 3.1 405B Instruct52.06#447
10Gemini 2.5 Flash51.41#237
11GPT-5 Mini49.58#176
12DeepSeek R1 052848.45#217
13GPT-4.1 Mini48.13#346
14DeepSeek V3 (0324)47#332
15Claude Sonnet 4 (20250514)46.18#211

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=kghalubench · How It Works · Data refreshed daily, snapshot 2026-10-11.