SEA-SpeechBench - Emotion Recognition (English Prompt): leaderboard

Metric: Judge-based accuracy (%; closed nine-class emotion label from speech, free-form answers mapped to labels and judged by Gemma-3-27B-Instruct, per-dataset scores averaged; short clips of at most 30 s from curated public Southeast Asian speech corpora, up to 1,000 sampled items per dataset; English instruction prompt). Source: arxiv.org. Saturation forecast: Around 2034. 15 models tracked.

Top models

#ModelScore
1Qwen2-Audio-7B-Instruct24.47
2Phi-4 Multimodal Instruct20.87
3Gemini 2.5 Flash19.5
4GPT-4o Audio17
5Qwen2.5-Omni-7B16.33
6Qwen3 Omni 30B A3B Instruct15.94
7Gemma 3n E4B (IT)12.46
8SeaLLMs-Audio-7B12.34
9gemma-3n-E2B-it12.21
10Voxtral-Mini-3B-250710.62

Interactive version: theaggregate.ai/benchmark?slug=sea-speechbench-emotion-recognition-english-prompt · How It Works · Data refreshed daily, snapshot 2026-09-26.