VocalAffectBench: leaderboard

Metric: Accuracy (%; seven-way expressed vocal emotion classification (angry, disgusted, fearful, happy, neutral, sad, surprised) of 280 human-recorded English clips, 40 per label, from raw audio alone with no transcript; prompted audio models get a fixed instruction listing the labels, provider emotion or prosody endpoints run at default settings, and native outputs are mapped to the label set before scoring). Source: arxiv.org. Saturation forecast: Around 2031. 6 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash44.3
2Qwen 3.5 Omni Plus37.9
3Voxtral-Small-24B-250734.3

Interactive version: theaggregate.ai/benchmark?slug=vocalaffectbench · How It Works · Data refreshed daily, snapshot 2026-09-26.