SEA-SpeechBench - Speech Translation (SEA Prompt): leaderboard

Metric: BLEU (0-100): corpus BLEU of the English translation of Southeast Asian speech; short clips of at most 30 s from curated public Southeast Asian speech corpora, up to 1,000 sampled items per dataset; instruction prompt in the native Southeast Asian language. Source: arxiv.org. Saturation forecast: Around 2028. 15 models tracked.

Top models

#ModelScore
1GPT-4o Audio21.39
2Gemini 2.5 Flash18.89
3Voxtral-Mini-3B-250718.15
4Gemma 3n E4B (IT)13.58
5Qwen3 Omni 30B A3B Instruct13.55
6SeaLLMs-Audio-7B10.11
7gemma-3n-E2B-it8.72
8Qwen2.5-Omni-7B8.06
9Qwen2-Audio-7B-Instruct3.2
10Phi-4 Multimodal Instruct0.32

Interactive version: theaggregate.ai/benchmark?slug=sea-speechbench-speech-translation-sea-prompt · How It Works · Data refreshed daily, snapshot 2026-09-26.