MOV-Bench (AOP-Agent): leaderboard
Metric: Accuracy (%; 519 audio-visual multi-hop multiple-choice questions; the AOP-Agent framework, hierarchical omni-modal memory and an observe-reflect-replan loop with every agent on the listed backbone). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen3-Omni-30B-A3B-Instruct (AOP-Agent) | 62.62 |
| 2 | Qwen3-Omni-30B-A3B-Thinking (AOP-Agent) | 62.24 |
| 3 | Qwen2.5-Omni-7B (AOP-Agent) | 47.01 |
Interactive version: theaggregate.ai/benchmark?slug=mov-bench-aop-agent · How It Works · Data refreshed daily, snapshot 2026-09-26.