HalluAudio - Speech: leaderboard
Metric: Average classification accuracy (%) over the ten speech task types of HalluAudio (human-verified audio question-answer pairs with adversarial, mixed-audio and unanswerable prompts; zero-shot; outputs normalized to labels, refusals counted wrong); unweighted mean of the task accuracies; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Phi-4 Multimodal Instruct | 35.47 |
| 2 | Qwen-Audio-Chat | 32.16 |
| 3 | Gemini 2.5 Flash | 16.6 |
Interactive version: theaggregate.ai/benchmark?slug=halluaudio-speech · How It Works · Data refreshed daily, snapshot 2026-10-07.