HalluScore - History: leaderboard

Metric: Factual hallucination rate (%) on the 240 historical HalluScore questions, zero-shot short-answer Arabic question answering at temperature 0 with no retrieval, responses annotated by humans; an answer that states an unsupported fact with confidence counts, an expressed uncertainty or refusal does not; lower is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.

Top models

#ModelScore
1GPT-529.17
2Claude Opus 437.5
3GPT-4o42.08
4Claude Sonnet 4.542.5
5O4 Mini52.92
6DeepSeek V355.42
7Llama 4 Maverick Instruct FP856.25
8Grok 4 Fast (Reasoning)57.5
9DeepSeek R161.67
10Qwen 3 Next 80B A3B Instruct67.92
11ALLaM-7B-Instruct-preview67.92
12Grok 470.83
13O383.33
14Fanar-1-9B84.17

Interactive version: theaggregate.ai/benchmark?slug=halluscore-history · How It Works · Data refreshed daily, snapshot 2026-10-07.