ODU-Bench (Audio-Only): leaderboard

Metric: Weighted demand-understanding score (%; 0.15 demand detection + 0.60 key-point coverage + 0.10 temporal localization + 0.10 transcription + 0.05 user profile, on 731 synthetic audio-only scenes (570 with a demand, 161 without), user turns as audio, earlier assistant replies as text; key points judged by Qwen3.6-Flash with thinking). Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1Seed 2.0 Lite77.3
2Gemini 3.1 Pro (Preview)75.4
3Qwen 3.5 Omni Plus74.7
4Gemini 3.7 Flash72.9
5Gemini 3.5 Flash Lite62.2
6Qwen3 Omni 30B A3B Instruct61.6
7MiniCPM-o-4.549.1
8Qwen2.5-Omni-7B48

Interactive version: theaggregate.ai/benchmark?slug=odu-bench-audio-only · How It Works · Data refreshed daily, snapshot 2026-09-26.