HaluEval-Wild: leaderboard
In-the-wild hallucination benchmark built from real user-chatbot interactions, adversarially filtered to measure hallucination rates on open-ended prompts.
Source: github.com.
Interactive version: theaggregate.ai/benchmark?slug=halueval-wild · How It Works · Data refreshed daily, snapshot 2026-09-05.