TraceAV-Bench - Visual-to-Audio Deception: leaderboard

Metric: Accuracy (%) on the 230 Visual-to-Audio Deception (V2A) questions of the multimodal hallucination dimension, four-option multiple choice over long audio-visual videos (10 to 140 minutes), exact match of the selected option set, each model at its official inference setting, with the full audio and video input; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)89.57
2Gemini 2.5 Pro79.13
3Gemini 3 Flash76.52
4Gemini 2.0 Flash74.78
5Gemma 4 E4B74.35
6Gemini 2.5 Flash60.87
7Qwen2.5-Omni-7B60.43

Interactive version: theaggregate.ai/benchmark?slug=traceav-bench-visual-to-audio-deception · How It Works · Data refreshed daily, snapshot 2026-10-07.