SEA-SpeechBench - Age Recognition (English Prompt): leaderboard

Metric: Macro-F1 (%; speaker age group (teens, adults, seniors) from speech, labels canonicalised by bilingual string matching, per-dataset scores averaged; short clips of at most 30 s from curated public Southeast Asian speech corpora, up to 1,000 sampled items per dataset; English instruction prompt). Source: arxiv.org. Saturation forecast: Around 2029. 15 models tracked.

Top models

#ModelScore
1Qwen3 Omni 30B A3B Instruct39.56
2Voxtral-Mini-3B-250737.26
3Gemini 2.5 Flash36.93
4GPT-4o Audio36.91
5Gemma 3n E4B (IT)32.41
6SeaLLMs-Audio-7B29.77
7gemma-3n-E2B-it26.82
8Phi-4 Multimodal Instruct26.28
9Qwen2-Audio-7B-Instruct17.23
10Qwen2.5-Omni-7B15.23

Interactive version: theaggregate.ai/benchmark?slug=sea-speechbench-age-recognition-english-prompt · How It Works · Data refreshed daily, snapshot 2026-09-26.