MSU-Bench - Background Inference: leaderboard
Metric: Accuracy (%) on the dialogue background inference task (Tier 2), four-option multiple choice on MSU-Bench (2,300 human-verified questions over Chinese and English multi-speaker conversations from telephone, meeting, podcast and movie audio), zero-shot with one instruction template requiring a single option letter, exact-match accuracy (x100); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash | 84 |
| 2 | Gemini 2.5 Flash | 81 |
| 3 | Gemini 2.5 Pro | 69 |
Interactive version: theaggregate.ai/benchmark?slug=msu-bench-background-inference · How It Works · Data refreshed daily, snapshot 2026-09-29.