HalluScore - Adversarial: leaderboard

Metric: Factual hallucination rate (%) on the 360 adversarial HalluScore questions (misleading or carefully crafted phrasing), zero-shot short-answer Arabic question answering at temperature 0 with no retrieval, responses annotated by humans; an answer that states an unsupported fact with confidence counts, an expressed uncertainty or refusal does not; lower is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.

Top models

#ModelScore
1GPT-531.11
2Claude Opus 434.72
3Claude Sonnet 4.538.89
4GPT-4o44.72
5O4 Mini51.11
6Grok 4 Fast (Reasoning)55.56
7Qwen 3 Next 80B A3B Instruct58.61
8DeepSeek R160.83
9DeepSeek V362.5
10Grok 465.28
11Llama 4 Maverick Instruct FP870.56
12ALLaM-7B-Instruct-preview75
13Fanar-1-9B80.83
14O381.67

Interactive version: theaggregate.ai/benchmark?slug=halluscore-adversarial · How It Works · Data refreshed daily, snapshot 2026-10-07.