SEA-SpeechBench - Gender Recognition (SEA Prompt): leaderboard

Metric: Macro-F1 (%; speaker gender from speech, labels canonicalised by bilingual string matching, per-dataset scores averaged; short clips of at most 30 s from curated public Southeast Asian speech corpora, up to 1,000 sampled items per dataset; instruction prompt in the native Southeast Asian language). Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash82.54
2Qwen2-Audio-7B-Instruct66.24
3Qwen3 Omni 30B A3B Instruct65.43
4Qwen2.5-Omni-7B49.17
5GPT-4o Audio46.78
6SeaLLMs-Audio-7B32.77
7Phi-4 Multimodal Instruct31.06
8Gemma 3n E4B (IT)23.54
9Voxtral-Mini-3B-250720.41

Interactive version: theaggregate.ai/benchmark?slug=sea-speechbench-gender-recognition-sea-prompt · How It Works · Data refreshed daily, snapshot 2026-09-26.