OmniStarPro-Long - CDQ (Medium Span): leaderboard
Metric: Recall accuracy (%; share of cross-event difference (CDQ) queries, which ask how a salient attribute such as a count, state or spatial arrangement changed between two widely separated moments, that a GPT-4o judge marks as matching the ground truth; evidence 10-30 minutes before the query; OmniStarPro-Long partition, streams of 10 minutes to over an hour). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | LiveStarPro | 42.8 |
| 2 | LiveStar | 28.1 |
| 3 | MMDuet | 19.2 |
| 4 | VideoLLM-MoD | 14.8 |
| 5 | VideoLLM-online | 14.1 |
Interactive version: theaggregate.ai/benchmark?slug=omnistarpro-long-cdq-medium-span · How It Works · Data refreshed daily, snapshot 2026-09-26.