AudioMC - Text Output: leaderboard

AudioMultiChallenge Text Output track benchmarks spoken dialogue systems that produce text responses across multi-turn interactions.

Metric: Score. Source: scale.com. Status: years away from saturation. 18 models tracked.

Top models

#ModelScore
1Inkling Small54.87
2Gemini 2.5 Pro (Thinking)46.9
3Gemini 2.5 Flash (Thinking)40.04
4Gemini 2.5 Flash26.11
5Gemma 3n E4B (IT)15.49
6Phi-4 Multimodal Instruct15.49

Interactive version: theaggregate.ai/benchmark?slug=audiomc-text-output · How It Works · Data refreshed daily, snapshot 2026-09-05.