SEA-SpeechBench - Speaker Recognition (English Prompt): leaderboard

Metric: Macro-F1 (%; whether two clips come from the same speaker, on self-constructed same and different speaker pairs, refusals count as errors; short clips of at most 30 s from curated public Southeast Asian speech corpora, up to 1,000 sampled items per dataset; English instruction prompt). Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash58.01
2Qwen3 Omni 30B A3B Instruct56.15
3SeaLLMs-Audio-7B43.5
4Voxtral-Mini-3B-250742.84
5Qwen2-Audio-7B-Instruct41.93
6gemma-3n-E2B-it37.16
7Phi-4 Multimodal Instruct36.43
8Gemma 3n E4B (IT)33.01

Interactive version: theaggregate.ai/benchmark?slug=sea-speechbench-speaker-recognition-english-prompt · How It Works · Data refreshed daily, snapshot 2026-09-26.