SEA-SpeechBench - Spoken QA (English Prompt): leaderboard

Metric: Judge score (0-100): short answers to questions about long-form Southeast Asian speech from YODAS2, rated 0-5 by a GPT-4.1 judge and multiplied by 20; English instruction prompt. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash92.28
2GPT-4o Audio86.66
3Qwen3 Omni 30B A3B Instruct85.77
4Voxtral-Mini-3B-250782.32
5Gemma 3n E4B (IT)79.24
6SeaLLMs-Audio-7B76.42
7gemma-3n-E2B-it73.35
8Qwen2.5-Omni-7B71.6
9Qwen2-Audio-7B-Instruct65.82
10Phi-4 Multimodal Instruct64.69

Interactive version: theaggregate.ai/benchmark?slug=sea-speechbench-spoken-qa-english-prompt · How It Works · Data refreshed daily, snapshot 2026-09-26.