MSU-Bench - Background Inference: leaderboard

Metric: Accuracy (%) on the dialogue background inference task (Tier 2), four-option multiple choice on MSU-Bench (2,300 human-verified questions over Chinese and English multi-speaker conversations from telephone, meeting, podcast and movie audio), zero-shot with one instruction template requiring a single option letter, exact-match accuracy (x100); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1Gemini 3 Flash84
2Gemini 2.5 Flash81
3Gemini 2.5 Pro69

Interactive version: theaggregate.ai/benchmark?slug=msu-bench-background-inference · How It Works · Data refreshed daily, snapshot 2026-09-29.