SEA-SpeechBench - Speech Translation (English Prompt): leaderboard

Metric: BLEU (0-100): corpus BLEU of the English translation of Southeast Asian speech; short clips of at most 30 s from curated public Southeast Asian speech corpora, up to 1,000 sampled items per dataset; English instruction prompt. Source: arxiv.org. Saturation forecast: Around 2029. 15 models tracked.

Top models

#ModelScore
1GPT-4o Audio21.24
2Voxtral-Mini-3B-250719.98
3Gemini 2.5 Flash16.86
4Qwen3 Omni 30B A3B Instruct15.28
5Gemma 3n E4B (IT)10.98
6SeaLLMs-Audio-7B10.74
7gemma-3n-E2B-it8.97
8Qwen2.5-Omni-7B7.91
9Qwen2-Audio-7B-Instruct4.54
10Phi-4 Multimodal Instruct3.04

Interactive version: theaggregate.ai/benchmark?slug=sea-speechbench-speech-translation-english-prompt · How It Works · Data refreshed daily, snapshot 2026-09-26.