VidOmni-Bench: leaderboard
Metric: F1 (%; sentence-wise event verification: for each of the dense captions written by five Video-LLMs for the 500 videos (4 s to 90 min, five complexity types), the model labels every sentence Correct or Incorrect and the human-labelled Incorrect sentences are the detection target; precision, recall and F1 are computed per video-caption pair and averaged over pairs; video-only input). Source: arxiv.org. Saturation forecast: Around May 2027. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash | 39.8 |
| 2 | GPT-5 Mini | 28.9 |
| 3 | Qwen 3 VL 8B | 14.7 |
| 4 | Keye-VL-1.5-8B | 9.6 |
Interactive version: theaggregate.ai/benchmark?slug=vidomni-bench · How It Works · Data refreshed daily, snapshot 2026-09-26.