SEA-SpeechBench - ASR: leaderboard
Metric: Word or character error rate (raw ratio, lower is better; WER, or CER for languages without explicit word boundaries, averaged over the source test sets of each language and then over the 11 languages (English, Filipino, Vietnamese, Indonesian, Tamil, Thai, Chinese, Khmer, Lao, Malay and Burmese); short clips of at most 30 s from curated public Southeast Asian speech corpora, up to 1,000 sampled items per dataset; English instruction prompt; exceeds 1 when a hypothesis is much longer than its reference). Source: arxiv.org. Saturation forecast: Around October 2026. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Flash | 0.15 |
| 2 | Qwen3 Omni 30B A3B Instruct | 0.56 |
| 3 | SeaLLMs-Audio-7B | 0.72 |
| 4 | Qwen2-Audio-7B-Instruct | 1 |
| 5 | Gemma 3n E4B (IT) | 1.14 |
| 6 | Qwen2.5-Omni-7B | 1.33 |
| 7 | Voxtral-Mini-3B-2507 | 1.36 |
| 8 | gemma-3n-E2B-it | 1.88 |
| 9 | Phi-4 Multimodal Instruct | 2.43 |
Interactive version: theaggregate.ai/benchmark?slug=sea-speechbench-asr · How It Works · Data refreshed daily, snapshot 2026-09-26.