ESCUCHA - Spoken Questions: leaderboard
Metric: Accuracy (%; 100 multiple-choice items whose question is delivered as synthesized speech inside the audio, options in text; in-the-wild Spanish audio from 6.7 s to 85 min per question with multiple accents and non-normative speech; one shared system prompt for every model). Source: arxiv.org. Saturation forecast: Around 2031. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen2.5-Omni-7B | 58 |
| 2 | Voxtral-Mini-3B-2507 | 55 |
| 3 | Gemini 2.5 Flash | 53 |
| 4 | Gemma 4 12B (IT) | 49 |
Interactive version: theaggregate.ai/benchmark?slug=escucha-spoken-questions · How It Works · Data refreshed daily, snapshot 2026-09-29.