AVI-Bench: leaderboard

Metric: Overall score (%): mean of the four stage averages (perception, understanding, reasoning, primitive sensation), on AVI-Bench (5,864 audio-visual items in 14 tasks over four cognitive stages), zero-shot with the same prompts for every omni-modal model; task scores normalized to percentages; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 26 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro (Preview 05-06)57.21
2Gemini 2.5 Flash (Preview 04-17)46.02
3Gemini 2.0 Flash44.97
4Qwen2.5-Omni-7B41.33
5Gemini 1.5 Pro40.38
6Gemini 1.5 Flash38.42
7Phi-4 Multimodal Instruct29.04

Interactive version: theaggregate.ai/benchmark?slug=avi-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.