TraceAV-Bench (Visual Only): leaderboard

Metric: Accuracy (%), macro-average of the 12 general sub-tasks, four-option multiple choice over long audio-visual videos (10 to 140 minutes), exact match of the selected option set, each model at its official inference setting, with visual input only: video frames without the audio track (vision-only MLLMs, and omni models with the audio stream removed); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.

Top models

#ModelScore
1Qwen3-VL-32B41.04
2Ming-Flash-Omni-2.038.55
3Qwen3-Omni-30B-A3B37.4
4Qwen3-VL-8B34.31

Interactive version: theaggregate.ai/benchmark?slug=traceav-bench-visual-only · How It Works · Data refreshed daily, snapshot 2026-10-07.