CASU - Contextual Reasoning: leaderboard

Metric: Accuracy (%; four-option multiple-choice questions on semi-synthetic auditory scenes that combine speech, sound events and background; contextual reasoning: reconcile a spoken claim with the events and background that support or contradict it). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1GPT-4o Audio74.02
2Gemini 2.0 Flash73.3
3Qwen3 Omni 30B A3B Instruct71.18
4Qwen2.5-Omni-7B62.06
5Voxtral-Small-24B-250752.52
6Qwen2-Audio-7B-Instruct44.02
7SALMONN-13B43.21

Interactive version: theaggregate.ai/benchmark?slug=casu-contextual-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-26.