HaluEval-Wild: leaderboard

In-the-wild hallucination benchmark built from real user-chatbot interactions, adversarially filtered to measure hallucination rates on open-ended prompts.

Source: github.com.

Interactive version: theaggregate.ai/benchmark?slug=halueval-wild · How It Works · Data refreshed daily, snapshot 2026-09-05.