MOV-Bench (AOP-Agent): leaderboard

Metric: Accuracy (%; 519 audio-visual multi-hop multiple-choice questions; the AOP-Agent framework, hierarchical omni-modal memory and an observe-reflect-replan loop with every agent on the listed backbone). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.

Top models

#ModelScore
1Qwen3-Omni-30B-A3B-Instruct (AOP-Agent)62.62
2Qwen3-Omni-30B-A3B-Thinking (AOP-Agent)62.24
3Qwen2.5-Omni-7B (AOP-Agent)47.01

Interactive version: theaggregate.ai/benchmark?slug=mov-bench-aop-agent · How It Works · Data refreshed daily, snapshot 2026-09-26.