RAGTruth: leaderboard

Retrieval-augmented generation hallucination corpus for evaluating grounded factuality and hallucination detection in RAG outputs.

Metric: RAG Truth (self-reported). Source: benchmarklist.com. Status: saturated. 11 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet86.1
2Mistral Large 2 (Jul)85.9
3GPT-4o (2024-05-13)84.3
4Llama 3.1 405B Instruct82.9
5Llama 3.3 70B Instruct82.6
6Qwen 2.5 72B Instruct81.9
7QwQ 32B-Preview72.4

Interactive version: theaggregate.ai/benchmark?slug=ragtruth · How It Works · Data refreshed daily, snapshot 2026-09-05.