MSI-Bench - English Background Speech Retrieval: leaderboard
Metric: All-pass rate (%; the 96 background-speech-retrieval cases, where the decisive facts occur only in overlapping background speech, English split; a case passes when all its atomic rubrics pass, judged by DeepSeek V4 Pro on the dialogue transcript, with a deterministic tool-call validator as one more rubric for cases that need a tool call; one prediction per model and case from the audio of a multi-party, multi-turn scene ending in an assistant-directed handoff; open-weight models served with JSON-schema constrained decoding). Source: arxiv.org. Saturation forecast: Around 2028. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 19.8 |
| 2 | GPT Audio | 14.6 |
| 3 | Gemini 3.5 Flash | 9.4 |
| 4 | GPT Audio Mini | 7.3 |
| 5 | Qwen2.5-Omni-7B | 2.1 |
| 6 | Gemma 4 12B (Reasoning) | 2.1 |
| 7 | Gemma 4 12B (Non-reasoning) | 2.1 |
| 8 | Phi-4 Multimodal Instruct | 0 |
Interactive version: theaggregate.ai/benchmark?slug=msi-bench-english-background-speech-retrieval · How It Works · Data refreshed daily, snapshot 2026-09-26.