HaluEval-Wild — leaderboard
In-the-wild hallucination benchmark built from real user-chatbot interactions, adversarially filtered to measure hallucination rates on challenging open-ended prompts.
Source: github.com.
Interactive version: theaggregate.ai/benchmark?slug=halueval-wild · How the rankings work · Data refreshed daily, snapshot 2026-07-22.