MSI-Bench - English Selective Disclosure: leaderboard

Metric: All-pass rate (%; the 96 selective-disclosure cases, where earlier speech says which facts to withhold from which listeners and those listeners later ask for them, English split; a case passes when all its atomic rubrics pass, judged by DeepSeek V4 Pro on the dialogue transcript, with a deterministic tool-call validator as one more rubric for cases that need a tool call; one prediction per model and case from the audio of a multi-party, multi-turn scene ending in an assistant-directed handoff; open-weight models served with JSON-schema constrained decoding). Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)85.4
2Gemini 3.5 Flash67.7
3GPT Audio54.2
4Qwen2.5-Omni-7B43.8
5Gemma 4 12B (Non-reasoning)38.5
6Gemma 4 12B (Reasoning)33.3
7Phi-4 Multimodal Instruct32.3
8GPT Audio Mini11.5

Interactive version: theaggregate.ai/benchmark?slug=msi-bench-english-selective-disclosure · How It Works · Data refreshed daily, snapshot 2026-09-26.