HaluEval — leaderboard
Large-scale hallucination evaluation benchmark covering QA, dialogue, and summarization. Tests whether models can recognize generated content that conflicts with source material or factual knowledge.
Source: github.com.
Interactive version: theaggregate.ai/benchmark?slug=halueval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.