TraceAV-Bench - Audio-to-Visual Deception: leaderboard

Metric: Accuracy (%) on the 229 Audio-to-Visual Deception (A2V) questions of the multimodal hallucination dimension, four-option multiple choice over long audio-visual videos (10 to 140 minutes), exact match of the selected option set, each model at its official inference setting, with the full audio and video input; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Gemini 2.0 Flash81.66
2Gemini 3.1 Pro (Preview)79.91
3Gemini 3 Flash75.55
4Gemini 2.5 Pro74.24
5Gemma 4 E4B69.43
6Gemini 2.5 Flash66.81
7Qwen2.5-Omni-7B55.46

Interactive version: theaggregate.ai/benchmark?slug=traceav-bench-audio-to-visual-deception · How It Works · Data refreshed daily, snapshot 2026-10-07.