HalluAudio - Speech: leaderboard

Metric: Average classification accuracy (%) over the ten speech task types of HalluAudio (human-verified audio question-answer pairs with adversarial, mixed-audio and unanswerable prompts; zero-shot; outputs normalized to labels, refusals counted wrong); unweighted mean of the task accuracies; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 11 models tracked.

Top models

#ModelScore
1Phi-4 Multimodal Instruct35.47
2Qwen-Audio-Chat32.16
3Gemini 2.5 Flash16.6

Interactive version: theaggregate.ai/benchmark?slug=halluaudio-speech · How It Works · Data refreshed daily, snapshot 2026-10-07.