VidOmni-Bench: leaderboard

Metric: F1 (%; sentence-wise event verification: for each of the dense captions written by five Video-LLMs for the 500 videos (4 s to 90 min, five complexity types), the model labels every sentence Correct or Incorrect and the human-labelled Incorrect sentences are the detection target; precision, recall and F1 are computed per video-caption pair and averaged over pairs; video-only input). Source: arxiv.org. Saturation forecast: Around May 2027. 13 models tracked.

Top models

#ModelScore
1Gemini 3 Flash39.8
2GPT-5 Mini28.9
3Qwen 3 VL 8B14.7
4Keye-VL-1.5-8B9.6

Interactive version: theaggregate.ai/benchmark?slug=vidomni-bench · How It Works · Data refreshed daily, snapshot 2026-09-26.