AuthorityBench - Fabricated Citations on True Claims: leaderboard

Metric: Hallucination rate (%) on true claims presented with a fabricated citation (TC x FC condition): share of non-refused responses a Qwen3-8B judge labels hallucinated (the model denies or distorts the true claim); 15K stratified subset for four models, full dataset for Gemma 3 4B, Llama 3.1 8B Instruct and Phi-4 Mini Instruct; lower is better. Source: arxiv.org. Saturation forecast: Around November 2027. 7 models tracked.

Top models

#ModelScore
1DeepSeek V3.216.28
2Gemma 4 31B22.73
3Gemma 3 4B24.14
4GPT-5.4 Mini37.27
5Phi-4 Mini Instruct40.63
6Llama 3.1 8B Instruct43.76
7Claude Haiku 4.547.11

Interactive version: theaggregate.ai/benchmark?slug=authoritybench-fabricated-citations-on-true-claims · How It Works · Data refreshed daily, snapshot 2026-09-29.