SEA-SpeechBench - Spoken QA (SEA Prompt): leaderboard

Metric: Judge score (0-100): short answers to questions about long-form Southeast Asian speech from YODAS2, rated 0-5 by a GPT-4.1 judge and multiplied by 20; instruction prompt in the native Southeast Asian language. Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash86.21
2Qwen3 Omni 30B A3B Instruct85.62
3GPT-4o Audio82.44
4Voxtral-Mini-3B-250779.06
5Gemma 3n E4B (IT)78.89
6SeaLLMs-Audio-7B78.38
7Qwen2.5-Omni-7B75.57
8Qwen2-Audio-7B-Instruct57.95
9gemma-3n-E2B-it56.06
10Phi-4 Multimodal Instruct47.36

Interactive version: theaggregate.ai/benchmark?slug=sea-speechbench-spoken-qa-sea-prompt · How It Works · Data refreshed daily, snapshot 2026-09-26.