HalluScore - Factual Hallucination: leaderboard

Metric: Factual hallucination rate (%): share of the 827 HalluScore questions (hallucination-prone Modern Standard Arabic questions across eleven types and many domains), zero-shot short-answer Arabic question answering at temperature 0 with no retrieval, responses annotated by humans; an answer that states an unsupported fact with confidence counts, an expressed uncertainty or refusal does not; lower is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.

Top models

#ModelScore
1GPT-525.15
2Claude Opus 433.01
3Claude Sonnet 4.535.79
4GPT-4o41.23
5O4 Mini46.43
6Grok 4 Fast (Reasoning)52.96
7DeepSeek V354.78
8Llama 4 Maverick Instruct FP857.56
9DeepSeek R157.68
10Qwen 3 Next 80B A3B Instruct61.43
11Grok 465.42
12ALLaM-7B-Instruct-preview68.68
13Fanar-1-9B79.32
14O380.05

Interactive version: theaggregate.ai/benchmark?slug=halluscore-factual-hallucination · How It Works · Data refreshed daily, snapshot 2026-10-07.