OmniStarPro-Long - CDQ (Medium Span): leaderboard

Metric: Recall accuracy (%; share of cross-event difference (CDQ) queries, which ask how a salient attribute such as a count, state or spatial arrangement changed between two widely separated moments, that a GPT-4o judge marks as matching the ground truth; evidence 10-30 minutes before the query; OmniStarPro-Long partition, streams of 10 minutes to over an hour). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.

Top models

#ModelScore
1LiveStarPro42.8
2LiveStar28.1
3MMDuet19.2
4VideoLLM-MoD14.8
5VideoLLM-online14.1

Interactive version: theaggregate.ai/benchmark?slug=omnistarpro-long-cdq-medium-span · How It Works · Data refreshed daily, snapshot 2026-09-26.