SEA-SpeechBench - Timestamped Content Query (30-60 s): leaderboard

Metric: WER or CER (%; transcribe only the speech inside a queried time window of a long recording of 30-60 seconds, CER for languages without word boundaries; English instruction prompt; blank where the model cannot take audio of that length; can exceed 100; lower is better). Source: arxiv.org. Saturation forecast: Around October 2026. 13 models tracked.

Top models

#ModelScore
1Qwen3 Omni 30B A3B Instruct1.22
2Gemini 2.5 Flash2.64
3Qwen2.5-Omni-7B6.58
4GPT-4o Audio6.82
5Gemma 3n E4B (IT)7.38
6gemma-3n-E2B-it7.49
7Voxtral-Mini-3B-25077.74
8Phi-4 Multimodal Instruct14.9

Interactive version: theaggregate.ai/benchmark?slug=sea-speechbench-timestamped-content-query-30-60-s · How It Works · Data refreshed daily, snapshot 2026-09-26.