Vectara Hallucination Leaderboard — leaderboard
Measures how much LLMs hallucinate when summarizing text. Uses the Hughes Hallucination Evaluation Model (HHEM) for automated detection.
Metric: Factual Consistency Rate (%). Source: github.com. 105 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 Nano | 96.9 |
| 2 | Gemini 2.5 Flash Lite | 96.7 |
| 3 | Phi-4 | 96.3 |
| 4 | Llama 3.3 70B Instruct | 95.9 |
| 5 | Gemma 3 12B (IT) | 95.6 |
| 6 | Mistral Large 2 (Nov) Instruct (2411) | 95.5 |
| 7 | Qwen 3 8B | 95.2 |
| 8 | Nova Pro (v1) | 94.9 |
| 9 | nova-2-lite-v1 | 94.9 |
| 10 | Mistral Small 3 | 94.9 |
| 11 | Gemma 4 26B A4B (IT) | 94.8 |
| 12 | Granite 4.0 H Small | 94.8 |
| 13 | DeepSeek V3.2 Exp | 94.7 |
| 14 | Qwen 3 14B | 94.6 |
| 15 | DeepSeek V3.1 | 94.5 |
Interactive version: theaggregate.ai/benchmark?slug=vectara-hallucination-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.