OmniPro: leaderboard
Metric: Probe-mode accuracy (%): for each ground-truth trigger the model, given the video so far, is queried 2-5 s before (expecting a negative answer) and 0-3 s after (expecting the task answer, exact match on a structured output), and the trigger counts only if both are right, mean of the nine sub-tasks of OmniPro (2,700 human-verified proactive streaming questions on LongVALE and COIN videos, 84% needing audio; 1 fps input); higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen2.5-Omni-7B | 20.1 |
| 2 | Phi-4 Multimodal Instruct | 12.9 |
Interactive version: theaggregate.ai/benchmark?slug=omnipro · How It Works · Data refreshed daily, snapshot 2026-10-07.