OmniStarPro-FDQ (Online): leaderboard
Metric: Semantic correctness (0-10; GPT-4o judge score of answers to frame-level dense questions asked at dense intervals about the current frame, the model timing its own answers). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | LiveStarPro | 6.61 |
| 2 | LiveStar | 6.44 |
| 3 | MMDuet | 4.78 |
| 4 | VideoLLM-online | 2.35 |
| 5 | VideoLLM-MoD | 2.11 |
Interactive version: theaggregate.ai/benchmark?slug=omnistarpro-fdq-online · How It Works · Data refreshed daily, snapshot 2026-09-26.