VidOmni-Bench (Video + Audio): leaderboard
Metric: F1 (%; sentence-wise event verification: for each of the dense captions written by five Video-LLMs for the 500 videos (4 s to 90 min, five complexity types), the model labels every sentence Correct or Incorrect and the human-labelled Incorrect sentences are the detection target; precision, recall and F1 are computed per video-caption pair and averaged over pairs; video with its audio track). Source: arxiv.org. Saturation forecast: Around July 2027. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash | 41.7 |
Interactive version: theaggregate.ai/benchmark?slug=vidomni-bench-video-plus-audio · How It Works · Data refreshed daily, snapshot 2026-09-26.