SEA-SpeechBench - Emotion Recognition (SEA Prompt): leaderboard

Metric: Judge-based accuracy (%; closed nine-class emotion label from speech, free-form answers mapped to labels and judged by Gemma-3-27B-Instruct, per-dataset scores averaged; short clips of at most 30 s from curated public Southeast Asian speech corpora, up to 1,000 sampled items per dataset; instruction prompt in the native Southeast Asian language). Source: arxiv.org. Saturation forecast: Around 2030. 15 models tracked.

Top models

#ModelScore
1GPT-4o Audio19.5
2Qwen2-Audio-7B-Instruct19.36
3Gemini 2.5 Flash16.79
4Gemma 3n E4B (IT)13.93
5gemma-3n-E2B-it13.17
6Qwen3 Omni 30B A3B Instruct11.31
7Qwen2.5-Omni-7B10.45
8Phi-4 Multimodal Instruct9.2
9SeaLLMs-Audio-7B9.17
10Voxtral-Mini-3B-25075.35

Interactive version: theaggregate.ai/benchmark?slug=sea-speechbench-emotion-recognition-sea-prompt · How It Works · Data refreshed daily, snapshot 2026-09-26.