TraceAV-Bench (Audio Only): leaderboard
Metric: Accuracy (%), macro-average of the 12 general sub-tasks, four-option multiple choice over long audio-visual videos (10 to 140 minutes), exact match of the selected option set, each model at its official inference setting, with audio input only (audio LLMs, and omni models with the visual stream removed); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Ming-Flash-Omni-2.0 | 30.78 |
| 2 | Qwen2-Audio-7B | 30.74 |
| 3 | Qwen3-Omni-30B-A3B | 30.36 |
Interactive version: theaggregate.ai/benchmark?slug=traceav-bench-audio-only · How It Works · Data refreshed daily, snapshot 2026-10-07.