ESCUCHA - Spoken Questions: leaderboard

Metric: Accuracy (%; 100 multiple-choice items whose question is delivered as synthesized speech inside the audio, options in text; in-the-wild Spanish audio from 6.7 s to 85 min per question with multiple accents and non-normative speech; one shared system prompt for every model). Source: arxiv.org. Saturation forecast: Around 2031. 8 models tracked.

Top models

#ModelScore
1Qwen2.5-Omni-7B58
2Voxtral-Mini-3B-250755
3Gemini 2.5 Flash53
4Gemma 4 12B (IT)49

Interactive version: theaggregate.ai/benchmark?slug=escucha-spoken-questions · How It Works · Data refreshed daily, snapshot 2026-09-29.