SEA-SpeechBench - Speaker Recognition (SEA Prompt): leaderboard

Metric: Macro-F1 (%; whether two clips come from the same speaker, on self-constructed same and different speaker pairs, refusals count as errors; short clips of at most 30 s from curated public Southeast Asian speech corpora, up to 1,000 sampled items per dataset; instruction prompt in the native Southeast Asian language). Source: arxiv.org. Saturation forecast: Around January 2028. 15 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash51.54
2Qwen3 Omni 30B A3B Instruct44.28
3Voxtral-Mini-3B-250738.67
4Gemma 3n E4B (IT)34.13
5Phi-4 Multimodal Instruct32.73
6gemma-3n-E2B-it32.07
7Qwen2-Audio-7B-Instruct31.66
8SeaLLMs-Audio-7B31.05

Interactive version: theaggregate.ai/benchmark?slug=sea-speechbench-speaker-recognition-sea-prompt · How It Works · Data refreshed daily, snapshot 2026-09-26.