Vectara Hallucination Leaderboard — leaderboard

Measures how much LLMs hallucinate when summarizing text. Uses the Hughes Hallucination Evaluation Model (HHEM) for automated detection.

Metric: Factual Consistency Rate (%). Source: github.com. 105 models tracked.

Top models

#ModelScore
1GPT-5.4 Nano96.9
2Gemini 2.5 Flash Lite96.7
3Phi-496.3
4Llama 3.3 70B Instruct95.9
5Gemma 3 12B (IT)95.6
6Mistral Large 2 (Nov) Instruct (2411)95.5
7Qwen 3 8B95.2
8Nova Pro (v1)94.9
9nova-2-lite-v194.9
10Mistral Small 394.9
11Gemma 4 26B A4B (IT)94.8
12Granite 4.0 H Small94.8
13DeepSeek V3.2 Exp94.7
14Qwen 3 14B94.6
15DeepSeek V3.194.5

Interactive version: theaggregate.ai/benchmark?slug=vectara-hallucination-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.