MSU-Bench: leaderboard

Metric: Overall accuracy (%), instance-level over all 16 tasks, four-option multiple choice on MSU-Bench (2,300 human-verified questions over Chinese and English multi-speaker conversations from telephone, meeting, podcast and movie audio), zero-shot with one instruction template requiring a single option letter, exact-match accuracy (x100); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1Gemini 3 Flash77
2Gemini 2.5 Pro70
3Gemini 2.5 Flash69

Interactive version: theaggregate.ai/benchmark?slug=msu-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.