HalluScore - Reasoning: leaderboard

Metric: Factual hallucination rate (%) on the 220 HalluScore questions that require reasoning, zero-shot short-answer Arabic question answering at temperature 0 with no retrieval, responses annotated by humans; an answer that states an unsupported fact with confidence counts, an expressed uncertainty or refusal does not; lower is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.

Top models

#ModelScore
1GPT-520.91
2Claude Opus 426.82
3Claude Sonnet 4.530.91
4O4 Mini35.91
5Grok 4 Fast (Reasoning)42.73
6GPT-4o43.18
7DeepSeek V346.82
8DeepSeek R150.91
9Llama 4 Maverick Instruct FP851.36
10Qwen 3 Next 80B A3B Instruct60.91
11Grok 462.73
12ALLaM-7B-Instruct-preview68.64
13O377.73
14Fanar-1-9B80.91

Interactive version: theaggregate.ai/benchmark?slug=halluscore-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.