TraceAV-Bench - Speech Context: leaderboard

Metric: Accuracy (%) on the 130 Speech Context (SC) questions, four-option multiple choice over long audio-visual videos (10 to 140 minutes), exact match of the selected option set, each model at its official inference setting, with the full audio and video input; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)96.92
2Gemini 3 Flash86.92
3Gemini 2.5 Pro83.85
4Gemini 2.5 Flash81.54
5Gemini 2.0 Flash70
6Qwen2.5-Omni-7B60
7Gemma 4 E4B55.38

Interactive version: theaggregate.ai/benchmark?slug=traceav-bench-speech-context · How It Works · Data refreshed daily, snapshot 2026-10-07.