TraceAV-Bench (Visual Only): leaderboard
Metric: Accuracy (%), macro-average of the 12 general sub-tasks, four-option multiple choice over long audio-visual videos (10 to 140 minutes), exact match of the selected option set, each model at its official inference setting, with visual input only: video frames without the audio track (vision-only MLLMs, and omni models with the audio stream removed); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen3-VL-32B | 41.04 |
| 2 | Ming-Flash-Omni-2.0 | 38.55 |
| 3 | Qwen3-Omni-30B-A3B | 37.4 |
| 4 | Qwen3-VL-8B | 34.31 |
Interactive version: theaggregate.ai/benchmark?slug=traceav-bench-visual-only · How It Works · Data refreshed daily, snapshot 2026-10-07.